
A wheat breeder staring at a fragmented draft genome knows the frustration: the disease resistance gene cluster sits in a centromeric gap that short reads cannot span. A livestock geneticist calculating breeding values from unphased SNP data knows the limitation: two beneficial alleles on opposite haplotypes look like one mediocre signal. An evolutionary biologist comparing gene families across species knows the risk: collapsed paralogs in a short-read assembly produce spurious gene loss calls. These are not edge cases — they are the norm in animal and plant genomics, where large genomes, polyploidy, and repetitive architecture push short-read sequencing past its physical limits.
CD Genomics provides comprehensive long-read sequencing services for animal and plant genomics, powered by both PacBio HiFi and Oxford Nanopore Technologies (ONT) platforms. We offer nine specialized services — from de novo genome assembly and T2T gap-free genomes to pan-genome analysis and full-length transcriptomics — covering the complete research pipeline from sample to publication-ready data. We have sequenced hundreds of eukaryotic species, and if your species is not on our list, we will design a custom workflow for it. This page is your directory to all animal and plant genomics services within our LongSeq division.
At a glance:
Animal and plant genomes are structurally different from the compact, diploid genomes that made short-read sequencing famous. Wheat (Triticum aestivum) carries a 14.5 Gb hexaploid genome — three subgenomes, each with its own complement of transposable elements, nested within a single nucleus. Sugarcane (Saccharum officinarum) is octoploid. Cultivated strawberry (Fragaria × ananassa) is octoploid. Cotton (Gossypium hirsutum) is allotetraploid. Even diploid crop genomes like maize (Zea mays, ~2.3 Gb) are 85% repetitive, dominated by long terminal repeat retrotransposons that scatter across centromeres and intergenic space.
When a short-read sequencer fragments these genomes into 150 bp pieces and attempts computational reassembly, three things break. First, repeats longer than the read length collapse — transposable elements, segmental duplications, and centromeric satellites disappear into single collapsed contigs. Second, haplotypes blend — in a polyploid, it becomes impossible to assign a given read to its subgenome of origin, turning three distinct subgenomes into one consensus blur. Third, structural variants vanish — insertions, deletions, inversions, and copy-number variations larger than the insert size are invisible to the alignment algorithm. A 2025 study that assembled the first T2T hexaploid wheat genome found that short-read assemblies missed entire gene clusters embedded in centromeric and pericentromeric regions — precisely the regions that harbor agronomically important disease resistance loci.
Long-read sequencing — with reads routinely exceeding 20 kb for PacBio HiFi and 100 kb for Oxford Nanopore — spans these repetitive regions natively. A single PacBio HiFi read covers a full-length LTR retrotransposon and its flanking unique sequence. An ONT ultra-long read traverses an entire centromeric satellite array. The result is a contiguous, phased assembly that preserves subgenome identity, captures structural variation, and reveals the complete gene repertoire — not just the fraction that happens to sit in unique sequence.
CD Genomics operates both major long-read platforms — PacBio SMRT sequencing and Oxford Nanopore Technologies — giving you access to the right technology for your specific genome and research question. This dual-platform capability matters because plant and animal genomes present different challenges, and no single platform is universally optimal.
PacBio HiFi sequencing generates highly accurate (Q30+, >99.9% consensus accuracy) reads of 15-25 kb. For projects requiring a high-quality reference genome — especially polyploid crops where subgenome phasing accuracy is critical, or livestock genomes destined for variant discovery and breeding value estimation — HiFi reads provide the accuracy and contiguity needed for chromosome-scale assembly. Paired with Hi-C chromatin conformation capture, a PacBio HiFi assembly routinely achieves chromosome-arm or chromosome-level scaffolds, with BUSCO completeness scores exceeding 95% for embryophyte and vertebrate lineages. Visit our PacBio SMRT Sequencing Technology hub for detailed platform information.
Oxford Nanopore sequencing generates reads up to megabases in length — the longest commercially available. For T2T genome assemblies that must span centromeric satellite arrays, telomeric repeats, and ribosomal DNA clusters, ONT ultra-long reads provide the physical connectivity that bridges these extreme repetitive regions. ONT also enables direct RNA sequencing — reading native RNA molecules without reverse transcription or amplification — which preserves base modification information (m6A, pseudouridine) and eliminates PCR bias. For animal and plant transcriptomics in non-model species, this is transformative. See our Oxford Nanopore Sequencing Technology hub for platform details.
At CD Genomics, our project consultation team helps you select the right platform — or combination of platforms — based on your species, genome characteristics, and research goals. We do not push a single technology; we design the workflow that best answers your question.
Below is our complete animal and plant genomics service catalog. Each linked service page includes detailed workflows, sample requirements, bioinformatics deliverables, and project consultation information.
Whether you are building a reference genome from scratch or characterizing population-level variation against an existing assembly, long-read sequencing delivers the contiguity and completeness that short reads cannot match.
For services not listed here, our sibling category hubs cover Human Genomics with Long-Read Sequencing, Microbial Genomics with Long-Read Sequencing, and Transcriptomics with Long-Read Sequencing.
The species listed below represent our highest-volume projects. This list demonstrates the breadth of our experience across taxonomic groups, genome architectures, and research applications.
Fundamental biological research relies on well-characterized model systems. CD Genomics has supported genomic studies in mouse (Mus musculus), rat (Rattus norvegicus), zebrafish (Danio rerio), fruit fly (Drosophila melanogaster), nematode (Caenorhabditis elegans), Arabidopsis thaliana, and rice (Oryza sativa). For model organism researchers transitioning from short-read to long-read platforms, we provide comparative data showing the additional genomic features captured by long reads.
Agricultural genomics and veterinary research represent a major fraction of our animal projects. Commonly sequenced livestock species include cattle (Bos taurus / B. indicus), pig (Sus scrofa), sheep (Ovis aries), goat (Capra hircus), chicken (Gallus gallus), horse (Equus caballus), duck (Anas platyrhynchos), and rabbit (Oryctolagus cuniculus). Companion animal projects include dog (Canis familiaris) and cat (Felis catus).
Aquaculture is the fastest-growing food production sector globally, and genomics is central to genetic improvement programs. We have sequenced Atlantic salmon (Salmo salar) — a species with a pseudo-tetraploid genome resulting from an ancestral whole-genome duplication — rainbow trout (Oncorhynchus mykiss), tilapia (Oreochromis niloticus), common carp (Cyprinus carpio), channel catfish (Ictalurus punctatus), grass carp (Ctenopharyngodon idella), Pacific white shrimp (Litopenaeus vannamei), Pacific oyster (Crassostrea gigas), and sea bass (Dicentrarchus labrax). Long reads are particularly important for aquaculture species because many lack reference genomes, and their genomes often contain high repeat content and residual polyploidy.
Plant genome projects demand long-read sequencing because of polyploidy, large genome size, and repeat content. Our most frequently sequenced crop and horticultural species include:
For polyploid species, we explicitly note ploidy level because it determines assembly strategy — a hexaploid wheat genome requires subgenome-aware assembly algorithms that differ fundamentally from those used for diploid rice.
Conservation genomics often involves limited or degraded samples from non-model species. CD Genomics has experience with giant panda (Ailuropoda melanoleuca), tiger (Panthera tigris), various avian, reptilian, amphibian, and insect species, as well as marine mammals. We offer low-input DNA extraction protocols optimized for non-invasive samples such as feathers, hair, scat, and shed skin.
The species catalog above represents our most frequent projects — but it is not a menu. CD Genomics' long-read sequencing platforms and bioinformatics pipelines are species-agnostic. Our assembly algorithms do not require a reference genome. Our annotation pipelines can be configured for any eukaryotic lineage using BUSCO lineage datasets, repeat libraries, and ab initio gene predictors appropriate to your species' taxonomic group.
We have successfully delivered genome assemblies, transcriptomes, and comparative analyses for insects (Lepidoptera, Coleoptera, Hymenoptera), reptiles (snakes, lizards, turtles), amphibians (frogs, salamanders), fungi (Ascomycota, Basidiomycota), algae (Chlorophyta, Rhodophyta), forest trees (poplar, eucalyptus, oak, pine, spruce), and marine invertebrates (corals, sea urchins, mollusks). Each project begins with a consultation: you tell us your species, your genome size estimate (if known), and your research question, and our team designs a customized workflow with platform recommendation, sequencing depth calculation, and bioinformatics plan.
How to start: Contact our project consultation team through the inquiry form with your species name, estimated genome size, ploidy level (if known), sample type and quantity, and your primary research question. We will respond with a customized project proposal within 1-2 business days.
Choosing between PacBio HiFi and Oxford Nanopore for animal and plant genomics depends on your genome characteristics and research goals. Below is a decision-oriented comparison.
| Feature | PacBio HiFi | Oxford Nanopore (ONT) |
| Read Length | 15-25 kb (HiFi) | 50 kb - 1+ Mb (ultra-long) |
| Consensus Accuracy | Q30+ (>99.9%) | Q20+ (duplex); Q30+ with polishing |
| Best for Polyploid Genomes | Excellent — high accuracy supports subgenome phasing with Hifiasm trio-binning | Good — ultra-long reads span large haplotype blocks, accuracy requires polishing |
| Best for T2T Assembly | Good for euchromatic regions; needs ONT for centromeres and telomeres | Excellent — ultra-long reads span centromeric satellites and rDNA arrays |
| Organellar Genomes | Excellent — HiFi reads assemble mtDNA/cpDNA as single contigs | Good — long reads span organellar genomes; accuracy requires deeper coverage |
| Direct Modification Detection | 5mC via polymerase kinetics | 5mC, 5hmC, m6A, pseudouridine from nanopore signal |
| Cost for Large Genomes (>5 Gb) | Higher per-base cost | Lower per-base cost at scale |
| Recommended For | High-quality reference genomes; polyploid subgenome phasing; variant discovery at population scale | T2T assemblies; ultra-long SV detection; direct RNA sequencing |
Decision Guide: If you need the highest possible accuracy for a crop or livestock reference genome → start with PacBio HiFi, supplemented with ONT ultra-long reads if T2T completeness is required. If your primary goal is a T2T assembly with resolved centromeres → ONT ultra-long reads are essential, with PacBio HiFi for polishing. For non-model species needing both genome assembly and full-length transcript annotation → combine PacBio HiFi for the genome and ONT direct RNA-seq for the transcriptome. Most of our most successful projects combine both platforms, and our project team recommends the optimal strategy for your species.
For detailed platform information, visit our PacBio SMRT Sequencing Technology and Oxford Nanopore Sequencing Technology hub pages.
Animal and plant genomics projects present unique sample logistics. A crop geneticist collecting leaf tissue from a field trial faces different challenges than a marine biologist sampling fin clips on a research vessel. CD Genomics has processed samples from every continent and most major ecosystem types.
| Sample Type | Quantity | Quality Requirement | Special Handling |
| Fresh leaf tissue (plant) | 2-5 g | Young, healthy tissue; flash-frozen | High-polyphenol species require specialized CTAB/PVP extraction |
| Whole blood (livestock) | 2-5 mL | EDTA or ACD tube; non-coagulated | Ship on cold packs within 48 h |
| Animal tissue (muscle, liver) | 50-100 mg | Flash-frozen immediately post-collection | Avoid repeated freeze-thaw; RNAlater for RNA |
| Fin clip (fish) | ~50 mg | Ethanol-preserved (95%) or flash-frozen | Provide species name and collection location |
| Feather (bird) | 3-5 quill feathers | Plucked (not molted), intact calamus | Ship dry at room temperature in paper envelope |
| Fungal mycelium | 100-200 mg | Flash-frozen or lyophilized | Axenic culture preferred; provide growth conditions |
| Herbarium / museum | 20-100 mg | As preserved | Contact before shipping — specialized extraction for degraded DNA |
| Non-invasive (hair, scat) | Varies | As collected | Contact us for project-specific protocols |
For high-molecular-weight (HMW) DNA extraction — critical for long-read sequencing — we follow protocols optimized for plant and animal tissues. Plant HMW extraction uses nuclei isolation to separate DNA from organellar and polysaccharide contaminants. Animal HMW extraction uses gentle lysis with wide-bore pipette tips to preserve fragment length. Both protocols are validated on our sequencing platforms and include QC checkpoints: NanoDrop for purity, Qubit for concentration, and FEMTO Pulse for fragment size distribution.
Comprehensive sample preparation guidelines are available at our Sample Submission Guideline page.
Sequencing is the input — analysis is the output your publication depends on. Our bioinformatics team provides end-to-end data analysis for animal and plant genomics projects, using species-appropriate tools and reference datasets.
For de novo assembly, we deploy platform-appropriate assemblers: Hifiasm for PacBio HiFi data (with trio-binning mode for phased diploid/polyploid assemblies), Flye or Canu for ONT data, and Verkko for hybrid T2T assembly combining both data types. Assembly polishing uses Medaka (ONT), Racon, or DeepPolisher (HiFi). Quality assessment uses QUAST for contiguity metrics (N50, L50) and BUSCO with lineage-appropriate datasets: embryophyta_odb10 for plants, vertebrata_odb10 for vertebrates, arthropoda_odb10 for insects, fungi_odb10 for fungi.
Repeat masking uses RepeatModeler to build species-specific repeat libraries followed by RepeatMasker. Gene prediction combines ab initio methods (AUGUSTUS, BRAKER3 with RNA-seq evidence) and evidence-based approaches (MAKER pipeline integrating transcript and protein homology evidence). Functional annotation uses InterProScan, eggNOG-mapper, and BLAST against Swiss-Prot and NR databases. For plant genomes, we offer specialized annotation of resistance gene analogs, carbohydrate-active enzymes (dbCAN), and transcription factors.
For multi-species projects: orthology inference with OrthoFinder, phylogenomic tree construction with IQ-TREE or RAxML, gene family expansion/contraction analysis with CAFE, positive selection detection with PAML (codeml branch-site models), and synteny visualization with MCScanX and circos plots. For population genomics: structural variant calling with Sniffles2 or cuteSV, haplotype phasing with Whatshap, and GWAS incorporating structural variants alongside SNPs.
Full-length isoform discovery with FLAIR, Bambu, or TALON. Pan-genome construction with PanGenome Graph Builder (PGGB) or Minigraph-Cactus, including structural variant-based graph genomes and presence/absence variation calling.
Standard deliverables include: raw sequence data (FASTQ), aligned reads (BAM), assembled genome (FASTA), variant calls (VCF), annotation files (GFF3), and a comprehensive QC report. Advanced analysis deliverables are defined during project consultation.
Visit our Long-Read Sequencing Data Analysis Services hub for complete analysis service details.
Absolutely. Our assembly and annotation pipelines are designed to work without a reference genome. For transcriptomics, reference-free tools such as FLAIR and TALON perform isoform discovery directly from long reads. For genome assembly, Hifiasm, Flye, and Canu are de novo assemblers that do not require a reference. We routinely deliver high-quality genome assemblies for non-model species and customize annotation pipelines using BUSCO lineage datasets and ab initio gene predictors appropriate to your taxonomic group.
Polyploid genomes require subgenome-aware assembly strategies. For autopolyploids, we use Hifiasm with increased ploidy settings or phased assembly modes. For allopolyploids with known progenitor species, we use Hifiasm trio-binning or reference-guided phasing to separate subgenomes. For unknown ploidy levels, we first estimate ploidy from k-mer spectra (GenomeScope2) then select the appropriate strategy. Our experience with hexaploid wheat (T2T assembly spanning 14.51 Gb) and allotetraploid cotton provides a tested workflow for complex polyploid genomes.
The answer depends on your genome characteristics and research goals. PacBio HiFi provides the highest accuracy (Q30+) and is ideal for high-quality reference genomes, especially polyploid crops where phasing accuracy matters. ONT provides the longest reads (up to Mb scale), which are essential for T2T assemblies that must bridge centromeric and telomeric repeats. Many projects benefit from combining both: HiFi for base accuracy and subgenome phasing, ONT for gap closure across extreme repeats. Our project consultation team recommends the optimal strategy — we are platform-agnostic.
We accept fresh or flash-frozen tissue (leaf, muscle, liver), whole blood, fin clips, feathers, fungal mycelium, extracted DNA/RNA, and prepared libraries. For challenging samples — herbarium specimens, museum samples, non-invasive collection (hair, scat, shed skin) — we offer specialized extraction protocols. The key requirement for long-read sequencing is high-molecular-weight DNA: avoid vortexing, repeated freeze-thaw, and use wide-bore pipette tips. See our Sample Requirements table above and Sample Submission Guideline page for detailed protocols.
A standard de novo genome assembly project — from sample QC to chromosome-scale assembly and annotation — typically takes 8-12 weeks. Smaller genomes (<500 Mb) may complete in 6-8 weeks. Large polyploid genomes (>5 Gb) or T2T assemblies requiring multi-platform data may take 12-16 weeks. Targeted services such as amplicon sequencing or organellar genome assembly can complete in 3-4 weeks. Expedited timelines are available — discuss with your project manager during consultation.
All de novo genome assembly services include structural annotation (gene prediction) and functional annotation as standard deliverable. You receive the assembled genome (FASTA), gene models (GFF3), coding and protein sequences (CDS and PEP FASTA), functional annotation tables, repeat annotation, and non-coding RNA annotation. For plant genomes, we also annotate organellar genomes and provide complete chloroplast/mitochondrial sequences if they assemble from the total genomic DNA. Raw sequence data (FASTQ) and all intermediate files are also provided.
Yes. Comparative Genomic Analysis is one of our nine specialized services. We provide orthology inference, phylogenomic reconstruction, gene family expansion/contraction analysis, positive selection detection, synteny mapping, and genome alignment visualization. For projects comparing 3-50+ species, we deploy HPC pipelines that scale to the full eukaryotic tree of life. Complete, long-read-based assemblies are critical — fragmented short-read assemblies produce false gene duplication and loss calls that undermine evolutionary conclusions.
Contact our team through the inquiry form. To provide an accurate project proposal, we need: (1) your species name (scientific name if known), (2) estimated genome size (if unknown, we can estimate from k-mer analysis of a small sequencing run), (3) ploidy level (if known), (4) sample type, quantity, and collection/storage conditions, (5) your primary research question and desired deliverables, and (6) any timeline constraints. We respond with a customized project proposal including platform recommendation, sequencing strategy, bioinformatics plan, timeline, and quotation — typically within 1-2 business days.
1. Genome Assembly Continuity Comparison: Short-read assembly vs long-read assembly for a polyploid crop genome — contig N50, scaffold N50, and BUSCO completeness improvement with PacBio HiFi + ONT hybrid assembly.
2. BUSCO Completeness Across Species: Benchmarking Universal Single-Copy Ortholog scores across model organisms, livestock, crops, and non-model species assembled at CD Genomics, demonstrating consistent >95% completeness.
3. Polyploid Subgenome Phasing: Haplotype-resolved assembly visualization showing independent resolution of subgenomes in hexaploid wheat — three distinct haplotype blocks across a chromosome arm.

References

For research use only. Not for use in diagnostic procedures.