Animal and Plant Genomics with Long-Read Sequencing — PacBio & ONT Services

Animal and Plant Genomics with Long-Read Sequencing — PacBio & ONT Services

Animal and plant genomics long-read sequencing — DNA helix with tree-of-life showing animal, plant, and aquaculture species branches

A wheat breeder staring at a fragmented draft genome knows the frustration: the disease resistance gene cluster sits in a centromeric gap that short reads cannot span. A livestock geneticist calculating breeding values from unphased SNP data knows the limitation: two beneficial alleles on opposite haplotypes look like one mediocre signal. An evolutionary biologist comparing gene families across species knows the risk: collapsed paralogs in a short-read assembly produce spurious gene loss calls. These are not edge cases — they are the norm in animal and plant genomics, where large genomes, polyploidy, and repetitive architecture push short-read sequencing past its physical limits.

CD Genomics provides comprehensive long-read sequencing services for animal and plant genomics, powered by both PacBio HiFi and Oxford Nanopore Technologies (ONT) platforms. We offer nine specialized services — from de novo genome assembly and T2T gap-free genomes to pan-genome analysis and full-length transcriptomics — covering the complete research pipeline from sample to publication-ready data. We have sequenced hundreds of eukaryotic species, and if your species is not on our list, we will design a custom workflow for it. This page is your directory to all animal and plant genomics services within our LongSeq division.

Why Animal and Plant Genomes Demand Long-Read Sequencing

Animal and plant genomes are structurally different from the compact, diploid genomes that made short-read sequencing famous. Wheat (Triticum aestivum) carries a 14.5 Gb hexaploid genome — three subgenomes, each with its own complement of transposable elements, nested within a single nucleus. Sugarcane (Saccharum officinarum) is octoploid. Cultivated strawberry (Fragaria × ananassa) is octoploid. Cotton (Gossypium hirsutum) is allotetraploid. Even diploid crop genomes like maize (Zea mays, ~2.3 Gb) are 85% repetitive, dominated by long terminal repeat retrotransposons that scatter across centromeres and intergenic space.

When a short-read sequencer fragments these genomes into 150 bp pieces and attempts computational reassembly, three things break. First, repeats longer than the read length collapse — transposable elements, segmental duplications, and centromeric satellites disappear into single collapsed contigs. Second, haplotypes blend — in a polyploid, it becomes impossible to assign a given read to its subgenome of origin, turning three distinct subgenomes into one consensus blur. Third, structural variants vanish — insertions, deletions, inversions, and copy-number variations larger than the insert size are invisible to the alignment algorithm. A 2025 study that assembled the first T2T hexaploid wheat genome found that short-read assemblies missed entire gene clusters embedded in centromeric and pericentromeric regions — precisely the regions that harbor agronomically important disease resistance loci.

Long-read sequencing — with reads routinely exceeding 20 kb for PacBio HiFi and 100 kb for Oxford Nanopore — spans these repetitive regions natively. A single PacBio HiFi read covers a full-length LTR retrotransposon and its flanking unique sequence. An ONT ultra-long read traverses an entire centromeric satellite array. The result is a contiguous, phased assembly that preserves subgenome identity, captures structural variation, and reveals the complete gene repertoire — not just the fraction that happens to sit in unique sequence.

How Long-Read Sequencing Transforms Animal and Plant Genomics

CD Genomics operates both major long-read platforms — PacBio SMRT sequencing and Oxford Nanopore Technologies — giving you access to the right technology for your specific genome and research question. This dual-platform capability matters because plant and animal genomes present different challenges, and no single platform is universally optimal.

PacBio HiFi sequencing generates highly accurate (Q30+, >99.9% consensus accuracy) reads of 15-25 kb. For projects requiring a high-quality reference genome — especially polyploid crops where subgenome phasing accuracy is critical, or livestock genomes destined for variant discovery and breeding value estimation — HiFi reads provide the accuracy and contiguity needed for chromosome-scale assembly. Paired with Hi-C chromatin conformation capture, a PacBio HiFi assembly routinely achieves chromosome-arm or chromosome-level scaffolds, with BUSCO completeness scores exceeding 95% for embryophyte and vertebrate lineages. Visit our PacBio SMRT Sequencing Technology hub for detailed platform information.

Oxford Nanopore sequencing generates reads up to megabases in length — the longest commercially available. For T2T genome assemblies that must span centromeric satellite arrays, telomeric repeats, and ribosomal DNA clusters, ONT ultra-long reads provide the physical connectivity that bridges these extreme repetitive regions. ONT also enables direct RNA sequencing — reading native RNA molecules without reverse transcription or amplification — which preserves base modification information (m6A, pseudouridine) and eliminates PCR bias. For animal and plant transcriptomics in non-model species, this is transformative. See our Oxford Nanopore Sequencing Technology hub for platform details.

At CD Genomics, our project consultation team helps you select the right platform — or combination of platforms — based on your species, genome characteristics, and research goals. We do not push a single technology; we design the workflow that best answers your question.

Animal and Plant Genomics Services at CD Genomics

Below is our complete animal and plant genomics service catalog. Each linked service page includes detailed workflows, sample requirements, bioinformatics deliverables, and project consultation information.

Genome Assembly and Resequencing Services

Whether you are building a reference genome from scratch or characterizing population-level variation against an existing assembly, long-read sequencing delivers the contiguity and completeness that short reads cannot match.

Targeted Sequencing and Transcriptome Services

Advanced Genomic Analysis Services

For services not listed here, our sibling category hubs cover Human Genomics with Long-Read Sequencing, Microbial Genomics with Long-Read Sequencing, and Transcriptomics with Long-Read Sequencing.

Commonly Sequenced Species at CD Genomics

The species listed below represent our highest-volume projects. This list demonstrates the breadth of our experience across taxonomic groups, genome architectures, and research applications.

Model Organisms

Fundamental biological research relies on well-characterized model systems. CD Genomics has supported genomic studies in mouse (Mus musculus), rat (Rattus norvegicus), zebrafish (Danio rerio), fruit fly (Drosophila melanogaster), nematode (Caenorhabditis elegans), Arabidopsis thaliana, and rice (Oryza sativa). For model organism researchers transitioning from short-read to long-read platforms, we provide comparative data showing the additional genomic features captured by long reads.

Livestock and Companion Animals

Agricultural genomics and veterinary research represent a major fraction of our animal projects. Commonly sequenced livestock species include cattle (Bos taurus / B. indicus), pig (Sus scrofa), sheep (Ovis aries), goat (Capra hircus), chicken (Gallus gallus), horse (Equus caballus), duck (Anas platyrhynchos), and rabbit (Oryctolagus cuniculus). Companion animal projects include dog (Canis familiaris) and cat (Felis catus).

Aquaculture Species

Aquaculture is the fastest-growing food production sector globally, and genomics is central to genetic improvement programs. We have sequenced Atlantic salmon (Salmo salar) — a species with a pseudo-tetraploid genome resulting from an ancestral whole-genome duplication — rainbow trout (Oncorhynchus mykiss), tilapia (Oreochromis niloticus), common carp (Cyprinus carpio), channel catfish (Ictalurus punctatus), grass carp (Ctenopharyngodon idella), Pacific white shrimp (Litopenaeus vannamei), Pacific oyster (Crassostrea gigas), and sea bass (Dicentrarchus labrax). Long reads are particularly important for aquaculture species because many lack reference genomes, and their genomes often contain high repeat content and residual polyploidy.

Crop and Horticultural Plants

Plant genome projects demand long-read sequencing because of polyploidy, large genome size, and repeat content. Our most frequently sequenced crop and horticultural species include:

  • Cereals: wheat (Triticum aestivum, hexaploid, ~14.5 Gb), maize (Zea mays, ~2.3 Gb), rice (Oryza sativa, ~430 Mb), barley (Hordeum vulgare, ~5.3 Gb), sorghum (Sorghum bicolor)
  • Legumes: soybean (Glycine max, ~1.1 Gb), common bean (Phaseolus vulgaris), chickpea (Cicer arietinum), peanut (Arachis hypogaea, allotetraploid)
  • Oilseeds: rapeseed/canola (Brassica napus, allotetraploid), sunflower (Helianthus annuus), sesame (Sesamum indicum)
  • Fiber crops: cotton (Gossypium hirsutum, allotetraploid)
  • Vegetables: tomato (Solanum lycopersicum), potato (Solanum tuberosum, autotetraploid), pepper (Capsicum annuum), cucumber (Cucumis sativus)
  • Fruits: apple (Malus domestica), grape (Vitis vinifera), banana (Musa acuminata), citrus (Citrus sinensis), strawberry (Fragaria × ananassa, octoploid)
  • Sugar and stimulant crops: sugarcane (Saccharum officinarum, octoploid, ~10 Gb), coffee (Coffea arabica, allotetraploid), tea (Camellia sinensis)

For polyploid species, we explicitly note ploidy level because it determines assembly strategy — a hexaploid wheat genome requires subgenome-aware assembly algorithms that differ fundamentally from those used for diploid rice.

Wildlife and Conservation Species

Conservation genomics often involves limited or degraded samples from non-model species. CD Genomics has experience with giant panda (Ailuropoda melanoleuca), tiger (Panthera tigris), various avian, reptilian, amphibian, and insect species, as well as marine mammals. We offer low-input DNA extraction protocols optimized for non-invasive samples such as feathers, hair, scat, and shed skin.

Beyond the List — We Sequence Any Eukaryotic Species

The species catalog above represents our most frequent projects — but it is not a menu. CD Genomics' long-read sequencing platforms and bioinformatics pipelines are species-agnostic. Our assembly algorithms do not require a reference genome. Our annotation pipelines can be configured for any eukaryotic lineage using BUSCO lineage datasets, repeat libraries, and ab initio gene predictors appropriate to your species' taxonomic group.

We have successfully delivered genome assemblies, transcriptomes, and comparative analyses for insects (Lepidoptera, Coleoptera, Hymenoptera), reptiles (snakes, lizards, turtles), amphibians (frogs, salamanders), fungi (Ascomycota, Basidiomycota), algae (Chlorophyta, Rhodophyta), forest trees (poplar, eucalyptus, oak, pine, spruce), and marine invertebrates (corals, sea urchins, mollusks). Each project begins with a consultation: you tell us your species, your genome size estimate (if known), and your research question, and our team designs a customized workflow with platform recommendation, sequencing depth calculation, and bioinformatics plan.

How to start: Contact our project consultation team through the inquiry form with your species name, estimated genome size, ploidy level (if known), sample type and quantity, and your primary research question. We will respond with a customized project proposal within 1-2 business days.

Platform Selection for Animal and Plant Genomes

Choosing between PacBio HiFi and Oxford Nanopore for animal and plant genomics depends on your genome characteristics and research goals. Below is a decision-oriented comparison.

Feature PacBio HiFi Oxford Nanopore (ONT)
Read Length 15-25 kb (HiFi) 50 kb - 1+ Mb (ultra-long)
Consensus Accuracy Q30+ (>99.9%) Q20+ (duplex); Q30+ with polishing
Best for Polyploid Genomes Excellent — high accuracy supports subgenome phasing with Hifiasm trio-binning Good — ultra-long reads span large haplotype blocks, accuracy requires polishing
Best for T2T Assembly Good for euchromatic regions; needs ONT for centromeres and telomeres Excellent — ultra-long reads span centromeric satellites and rDNA arrays
Organellar Genomes Excellent — HiFi reads assemble mtDNA/cpDNA as single contigs Good — long reads span organellar genomes; accuracy requires deeper coverage
Direct Modification Detection 5mC via polymerase kinetics 5mC, 5hmC, m6A, pseudouridine from nanopore signal
Cost for Large Genomes (>5 Gb) Higher per-base cost Lower per-base cost at scale
Recommended For High-quality reference genomes; polyploid subgenome phasing; variant discovery at population scale T2T assemblies; ultra-long SV detection; direct RNA sequencing

Decision Guide: If you need the highest possible accuracy for a crop or livestock reference genome → start with PacBio HiFi, supplemented with ONT ultra-long reads if T2T completeness is required. If your primary goal is a T2T assembly with resolved centromeres → ONT ultra-long reads are essential, with PacBio HiFi for polishing. For non-model species needing both genome assembly and full-length transcript annotation → combine PacBio HiFi for the genome and ONT direct RNA-seq for the transcriptome. Most of our most successful projects combine both platforms, and our project team recommends the optimal strategy for your species.

For detailed platform information, visit our PacBio SMRT Sequencing Technology and Oxford Nanopore Sequencing Technology hub pages.

Sample Preparation and Requirements for Animal and Plant Genomics

Animal and plant genomics projects present unique sample logistics. A crop geneticist collecting leaf tissue from a field trial faces different challenges than a marine biologist sampling fin clips on a research vessel. CD Genomics has processed samples from every continent and most major ecosystem types.

Sample Type Quantity Quality Requirement Special Handling
Fresh leaf tissue (plant) 2-5 g Young, healthy tissue; flash-frozen High-polyphenol species require specialized CTAB/PVP extraction
Whole blood (livestock) 2-5 mL EDTA or ACD tube; non-coagulated Ship on cold packs within 48 h
Animal tissue (muscle, liver) 50-100 mg Flash-frozen immediately post-collection Avoid repeated freeze-thaw; RNAlater for RNA
Fin clip (fish) ~50 mg Ethanol-preserved (95%) or flash-frozen Provide species name and collection location
Feather (bird) 3-5 quill feathers Plucked (not molted), intact calamus Ship dry at room temperature in paper envelope
Fungal mycelium 100-200 mg Flash-frozen or lyophilized Axenic culture preferred; provide growth conditions
Herbarium / museum 20-100 mg As preserved Contact before shipping — specialized extraction for degraded DNA
Non-invasive (hair, scat) Varies As collected Contact us for project-specific protocols

For high-molecular-weight (HMW) DNA extraction — critical for long-read sequencing — we follow protocols optimized for plant and animal tissues. Plant HMW extraction uses nuclei isolation to separate DNA from organellar and polysaccharide contaminants. Animal HMW extraction uses gentle lysis with wide-bore pipette tips to preserve fragment length. Both protocols are validated on our sequencing platforms and include QC checkpoints: NanoDrop for purity, Qubit for concentration, and FEMTO Pulse for fragment size distribution.

Comprehensive sample preparation guidelines are available at our Sample Submission Guideline page.

Bioinformatics and Data Analysis for Animal and Plant Genomics

Sequencing is the input — analysis is the output your publication depends on. Our bioinformatics team provides end-to-end data analysis for animal and plant genomics projects, using species-appropriate tools and reference datasets.

Genome Assembly and Quality Assessment

For de novo assembly, we deploy platform-appropriate assemblers: Hifiasm for PacBio HiFi data (with trio-binning mode for phased diploid/polyploid assemblies), Flye or Canu for ONT data, and Verkko for hybrid T2T assembly combining both data types. Assembly polishing uses Medaka (ONT), Racon, or DeepPolisher (HiFi). Quality assessment uses QUAST for contiguity metrics (N50, L50) and BUSCO with lineage-appropriate datasets: embryophyta_odb10 for plants, vertebrata_odb10 for vertebrates, arthropoda_odb10 for insects, fungi_odb10 for fungi.

Genome Annotation

Repeat masking uses RepeatModeler to build species-specific repeat libraries followed by RepeatMasker. Gene prediction combines ab initio methods (AUGUSTUS, BRAKER3 with RNA-seq evidence) and evidence-based approaches (MAKER pipeline integrating transcript and protein homology evidence). Functional annotation uses InterProScan, eggNOG-mapper, and BLAST against Swiss-Prot and NR databases. For plant genomes, we offer specialized annotation of resistance gene analogs, carbohydrate-active enzymes (dbCAN), and transcription factors.

Comparative and Population Genomics

For multi-species projects: orthology inference with OrthoFinder, phylogenomic tree construction with IQ-TREE or RAxML, gene family expansion/contraction analysis with CAFE, positive selection detection with PAML (codeml branch-site models), and synteny visualization with MCScanX and circos plots. For population genomics: structural variant calling with Sniffles2 or cuteSV, haplotype phasing with Whatshap, and GWAS incorporating structural variants alongside SNPs.

Transcriptomics and Pan-Genomics

Full-length isoform discovery with FLAIR, Bambu, or TALON. Pan-genome construction with PanGenome Graph Builder (PGGB) or Minigraph-Cactus, including structural variant-based graph genomes and presence/absence variation calling.

Standard deliverables include: raw sequence data (FASTQ), aligned reads (BAM), assembled genome (FASTA), variant calls (VCF), annotation files (GFF3), and a comprehensive QC report. Advanced analysis deliverables are defined during project consultation.

Visit our Long-Read Sequencing Data Analysis Services hub for complete analysis service details.

Frequently Asked Questions

Demo

1. Genome Assembly Continuity Comparison: Short-read assembly vs long-read assembly for a polyploid crop genome — contig N50, scaffold N50, and BUSCO completeness improvement with PacBio HiFi + ONT hybrid assembly.

2. BUSCO Completeness Across Species: Benchmarking Universal Single-Copy Ortholog scores across model organisms, livestock, crops, and non-model species assembled at CD Genomics, demonstrating consistent >95% completeness.

3. Polyploid Subgenome Phasing: Haplotype-resolved assembly visualization showing independent resolution of subgenomes in hexaploid wheat — three distinct haplotype blocks across a chromosome arm.

Animal and plant genomics demo — genome assembly continuity comparison, BUSCO completeness, and polyploid subgenome phasing visualization

References

  1. Liu S, Li K, Dai X, et al. A telomere-to-telomere genome assembly coupled with multi-omic data provides insights into the evolution of hexaploid bread wheat. Nature Genetics, 57(4), 1008-1020 (2025). doi:10.1038/s41588-025-02137-x
  2. Luo J, Huang N, et al. Telomere-to-telomere genome assembly of a male pig provides insight into population structure and selection for body stature. Nature Genetics, 57, (2025). doi:10.1038/s41588-025-02433-6
  3. Plant pangenomes for crop improvement, biodiversity and evolution. Nature Reviews Genetics, 25, 304-320 (2024). doi:10.1038/s41576-024-00691-4

Animal and plant genomics — sequenced species organized by category: model organisms, livestock, aquaculture, crops, and conservation species

For research use only. Not for use in diagnostic procedures.

Get Your Instant Quote