
Animal and plant genomes present unique challenges — polyploidy, massive repeat content, high heterozygosity — that short-read sequencing cannot fully resolve. CD Genomics provides comprehensive long-read sequencing solutions on PacBio HiFi and Oxford Nanopore platforms, delivering T2T genome assemblies, haplotype-resolved phasing, and functional annotation for species ranging from crops and livestock to endangered wildlife. Whether your focus is agrigenomics innovation or biodiversity conservation, our end-to-end CRO services combine dual-platform flexibility with species-specific bioinformatics expertise.
At a glance:
Animal and plant genomes present structural challenges that short-read sequencing cannot fully overcome. Polyploidy in crops like wheat (allohexaploid, ~15 Gb) and sugarcane (autopolyploid, ~10 Gb) produces near-identical subgenomes that collapse into chimeric assemblies with short reads. High repeat content — often exceeding 70% in plant genomes — generates assembly gaps at transposable elements, centromeres, and segmental duplications. High heterozygosity in outbred livestock and wild populations fragments assemblies into separate haplotypes rather than producing a unified reference.
Long-read sequencing addresses these challenges directly. PacBio HiFi reads deliver 15–25 kb reads at >99.9% consensus accuracy, enabling phased diploid and polyploid assemblies through algorithms like hifiasm that distinguish haplotypes at the read level. Oxford Nanopore ultra-long reads routinely exceed 100 kb, with some reaching beyond 1 Mb, allowing individual reads to span entire repeat arrays, segmental duplications, and centromeric satellites — the very regions that short-read assemblies fail to resolve. When combined with Hi-C chromatin conformation capture, these technologies produce chromosome-scale, telomere-to-telomere (T2T) assemblies with contiguity metrics that were impossible just five years ago.
The practical difference is measurable. Where short-read assemblies of complex plant genomes routinely produce tens of thousands of contigs with N50 values in the tens of kilobases, long-read assemblies consistently achieve contig N50 in the megabase range. The sorghum T2T assembly completed by Li et al. in 2024 demonstrated this dramatically: it eliminated all 3,913 gaps present in the previous short-read reference genome, captured all 10 centromeres and 20 telomeres, and achieved a contig N50 of 71.1 Mb with 98.88% k-mer completeness. For animal genomes, similar advances have been achieved — from livestock species where haplotype-resolved assemblies enable precision breeding, to endangered wildlife where high-quality reference genomes support conservation genomics.
At CD Genomics, we provide long-read sequencing services for animal and plant research on both PacBio Revio and Oxford Nanopore PromethION platforms, combined with species-specific bioinformatics expertise. Whether your project involves a compact 700 Mb sorghum genome or a massive 15 Gb wheat genome, our team designs the optimal sequencing strategy for your species, ploidy level, and research objective.
Choosing the right sequencing platform for an animal or plant genome project depends on genome size, ploidy, repeat content, and research goals. The table below compares PacBio HiFi and Oxford Nanopore platforms specifically for non-human genome applications.
| Feature | PacBio HiFi (Revio) | Oxford Nanopore Ultra-Long (PromethION) |
| Read Length | 15–25 kb | 100 kb – 1+ Mb |
| Accuracy (Q-score) | >Q30 (99.9%) | ~Q20 (99%), improving with newer chemistry |
| Best For | Diploid/polyploid phased assembly, moderate-size genomes, organelle genomes | Ultra-large genomes (>3 Gb), spanning repeat arrays, T2T gap closure |
| Polyploid Phasing | Excellent — hifiasm produces fully phased diploid assemblies | Good — phasing improves with ultra-long reads but accuracy limitations require HiFi complement |
| DNA Modification Detection | 5mC via polymerase kinetics (indirect) | Direct detection of 5mC, 6mA, and other modifications without conversion |
| Throughput per Run | ~90 Gb (Revio SMRT Cell 25M) | ~200 Gb (PromethION flow cell) |
| Cost per Gb | Higher base cost; lower for moderate-coverage HiFi projects | Lower cost per Gb; higher coverage needed to compensate for lower per-read accuracy |
How we recommend platform selection: For polyploid crop genomes — wheat, sugarcane, strawberry, potato — we typically recommend a dual-platform strategy. PacBio HiFi reads provide the base-level accuracy needed for subgenome phasing with hifiasm, while ONT ultra-long reads bridge the remaining gaps at centromeres, telomeres, and large repeat arrays. For diploid animal genomes under 3 Gb with moderate repeat content, PacBio HiFi alone often produces near-complete assemblies when combined with Hi-C scaffolding. For very large genomes — conifers, amphibians, lungfish — ONT ultra-long reads are essential for spanning massive repeat expansions, and we supplement with PacBio HiFi for base-level accuracy.
For projects requiring DNA methylation profiling alongside genome assembly, ONT sequencing provides direct modification detection without the bisulfite conversion step that degrades DNA. This is particularly valuable for plant epigenomics, where methylation patterns regulate transposon silencing, development, and stress responses. For mitochondrial and chloroplast genome sequencing, PacBio HiFi's circular consensus mode routinely produces complete organelle genomes in single contigs.
Our project consultation includes a platform recommendation tailored to your species, genome size estimate, and research goals. We do not lock clients into one technology — we deploy the right combination of platforms for each project.
CD Genomics long-read sequencing workflow for animal and plant research: from sample collection to T2T genome assembly and functional annotation.
Global food production faces intensifying pressure from climate change, population growth, and emerging pathogens. Long-read sequencing provides the genomic resolution needed to accelerate crop improvement, livestock breeding, and agricultural biotechnology. CD Genomics supports agrigenomics programs with species-specific sequencing strategies and bioinformatics pipelines optimized for breeding-relevant analyses.
High-quality reference genomes and population-scale resequencing reveal structural variants, copy-number variations, and presence-absence variations that explain agronomic traits beyond SNP-level associations. We support breeding programs from reference genome construction through GWAS and genomic selection.
Precise characterization of transgene insertion sites, off-target effects, and genome editing outcomes requires long reads that span the full integration locus. PacBio HiFi and ONT sequencing verify on-target integration, detect unintended rearrangements, and confirm editing precision at single-nucleotide resolution.
Pathogen genomics and host-resistance gene discovery benefit from complete genome assemblies that capture effector gene clusters, resistance gene analogs, and mobile genetic elements often located in repeat-rich genomic regions inaccessible to short-read sequencing. LRS enables tracking pathogen evolution and identifying durable resistance loci.
Biodiversity loss driven by habitat destruction, climate change, and pollution demands genomic tools that can characterize genetic diversity, population structure, and adaptive potential across threatened species. Long-read sequencing delivers reference-quality genomes for non-model organisms — often from limited or degraded samples — enabling conservation genomics at unprecedented resolution.
Population genomics of wild and managed animal populations requires haplotype-resolved assemblies to detect inbreeding, deleterious alleles, and adaptive variants. LRS resolves structural variants and immune gene complexes (MHC, TLR) that are critical for population viability analyses but inaccessible to short-read approaches.
Understanding species' responses to climate change requires reference genomes that capture the genetic architecture of thermal tolerance, drought resistance, and phenological adaptation. LRS enables comparative genomics across populations and species distributed along environmental gradients, identifying candidate loci for climate adaptation.
Marine species — from teleost fish to cetaceans, mollusks to corals — often possess large, repeat-rich genomes shaped by ancient whole-genome duplications. LRS produces contiguous assemblies that resolve these evolutionary signatures, enabling phylogenomic reconstruction, adaptation studies, and fisheries management genomics.
Beyond our solutions-focused application tracks, CD Genomics provides a full portfolio of individual long-read sequencing services for animal and plant research. Each service can be ordered standalone or combined into an integrated project workflow.
High-coverage long-read resequencing for variant discovery — SNPs, indels, structural variants, and copy-number variations — with dual-platform flexibility and population-scale design options.
Complete de novo genome assembly for any species without a reference — from DNA extraction and library preparation through chromosome-scale assembly, polishing, and functional annotation.
Full-length amplification and sequencing of target genomic regions up to 20 kb — ideal for MHC haplotyping, resistance gene profiling, organelle genome validation, and targeted structural variant confirmation.
Full-length transcript sequencing without assembly — capture complete isoform structures, alternative splicing events, fusion transcripts, and allele-specific expression with PacBio Iso-Seq or ONT direct RNA sequencing.
Move beyond a single reference genome — construct species-level pan-genomes that capture core and variable gene content across multiple accessions, varieties, or populations to identify presence-absence variation driving phenotypic diversity.
Complete organelle genome assembly with long reads that resolve the full circular molecule — including complex repeat regions and structural rearrangements that fragment short-read organelle assemblies.
Genome-wide simple sequence repeat and short tandem repeat identification and genotyping — leveraging long reads to characterize repeat length, motif composition, and flanking sequence for marker development and genetic diversity assessment.
Phase both haplotypes independently to produce a diploid T2T assembly — critical for outbred species, hybrid crops, and any organism where allele-specific variation matters for trait mapping and functional genomics.
Gapless, telomere-to-telomere genome assembly resolving all chromosomes from one telomere to the other — the gold standard for reference genome construction in the post-short-read era, now achievable for diverse animal and plant species.
Our standard animal and plant genome analysis pipeline transforms raw sequencing data into publication-ready assemblies and annotations through a systematic workflow developed specifically for non-model and complex genomes.
Raw Data Processing and Quality Control. Sequencing reads undergo quality assessment with FastQC and NanoPlot, followed by adapter trimming and filtering. For PacBio HiFi data, we extract circular consensus sequences and filter by expected quality. For ONT data, we apply Guppy/Dorado basecalling with the latest models, followed by read quality filtering and length selection based on genome size estimates.
Genome Assembly. We deploy species-appropriate assembly algorithms: hifiasm for phased diploid and polyploid assembly from HiFi reads; Flye and NextDenovo for ONT-based assemblies with long-read error correction; and hybrid approaches combining both platforms. Hi-C chromatin conformation data is integrated with YaHS or 3D-DNA for chromosome-scale scaffolding. For organelle genomes, we use dedicated circular assembly pipelines that produce complete cpDNA/mtDNA molecules in single contigs.
Assembly Quality Assessment. Every assembly undergoes rigorous quality evaluation: contiguity metrics (N50, N90, L50, total assembly size vs. expected genome size), completeness via BUSCO against lineage-specific ortholog databases (embryophyta_odb10, metazoa_odb10), k-mer completeness and QV estimation with Merqury, and Hi-C contact map inspection for scaffolding accuracy. We report all metrics transparently and address any quality issues before annotation.
Genome Annotation. The annotation pipeline includes: repeat identification and masking with RepeatModeler and RepeatMasker using species-specific repeat libraries; protein-coding gene prediction with BRAKER3 integrating RNA-seq evidence and protein homology; functional annotation via BLAST against Swiss-Prot, InterProScan for domain identification, and KEGG pathway mapping. For non-model organisms, we build custom training sets to optimize gene prediction accuracy. The final annotation is delivered in GFF3 format with accompanying functional annotation tables.
Comparative and Population Genomics. For multi-sample projects, we offer variant calling (SNPs, indels, structural variants) with long-read-aware tools including Sniffles2 and cuteSV for SVs, DeepVariant or Clair3 for small variants, and population genetics analyses including nucleotide diversity (π), FST, Tajima's D, and selective sweep detection. Pan-genome construction uses iterative mapping and assembly or graph-based approaches depending on the number of accessions.
High-quality long-read sequencing begins with high-molecular-weight (HMW) DNA or intact RNA. The table below summarizes sample requirements for common animal and plant sample types. Specific protocols vary by species — our project team provides detailed collection and shipping guidelines during consultation.
| Sample Type | Recommended Quantity | Quality Requirement | Shipping Condition |
| Plant leaf tissue (fresh, young) | 2–5 g | Disease-free, rapidly frozen | Liquid nitrogen / dry ice |
| Plant seed / embryo | 1–2 g | Viable, surface-sterilized | Dry ice or silica gel (DNA) |
| Animal blood | 2–5 mL | EDTA tube, non-hemolyzed | Cold pack / dry ice |
| Animal muscle tissue | 100–200 mg | Fresh or flash-frozen | Dry ice |
| Animal liver / spleen | 50–100 mg | Flash-frozen within 30 min of collection | Dry ice |
| Insect (whole body) | 5–10 individuals | Ethanol-preserved or fresh-frozen | Dry ice |
| Fish fin clip / muscle | 50–100 mg | Ethanol-preserved or flash-frozen | Dry ice |
| Mollusk / crustacean tissue | 100–200 mg | Flash-frozen, avoid gut content | Dry ice |
| Fungal mycelium | 100–200 mg (wet weight) | Pure culture, washed | Dry ice |
| Cultured cells | 10⁶–10⁷ cells | Viability > 85%, washed in PBS | Dry ice |
HMW DNA Extraction Note: For PacBio HiFi and ONT ultra-long sequencing, we perform HMW DNA extraction using optimized protocols (CTAB for plants, phenol-chloroform or magnetic bead-based for animals) to maximize fragment length. DNA size distribution is verified by pulsed-field gel electrophoresis or Femto Pulse before library preparation. For challenging samples — herbarium specimens, formalin-fixed tissues, or environmental samples with degraded DNA — we offer specialized extraction protocols and can advise on feasibility during consultation.
RNA Sample Note: For RNA sequencing services, we recommend flash-freezing tissues immediately upon collection and storing at −80°C. RNA integrity (RIN ≥ 7.0 for most applications; RIN ≥ 8.0 for Iso-Seq) is verified before library preparation. RNAlater stabilization is acceptable but may reduce polyA+ RNA yield — we recommend consultation before using preservatives.
Dean LL, Holmes N, Dobbs P, Loose M. The tiger who came to T2T: Telomere-to-Telomere genome assembly of the Sumatran tiger (Panthera tigris sumatrae) using nanopore simplex reads. BMC Genomics. 2026;27:339. (CC BY 4.0)
The Sumatran tiger (Panthera tigris sumatrae) is critically endangered, with fewer than 500 individuals estimated in the wild. High-quality reference genomes are essential for conservation genomics — they enable population genetic monitoring, identification of deleterious alleles, and informed management of captive breeding programs. However, producing chromosome-scale assemblies for non-model vertebrates has traditionally required multiple sequencing platforms (PacBio + ONT + Illumina + Hi-C), making projects expensive and technically complex. This study asked whether a single long-read technology — Oxford Nanopore simplex reads with computational error correction — could produce a near-T2T assembly at substantially reduced cost and complexity.
The researchers generated ONT simplex long-read data from a Sumatran tiger sample and evaluated three error-correction strategies: NextDenovo, HERRO, and hifiasm (ONT mode). Corrected reads were assembled with hifiasm and scaffolded against a previously published tiger reference genome. Assembly quality was assessed by contiguity (N50, longest contig), telomere repeat identification, and BUSCO completeness. Synteny comparisons were performed against the domestic cat genome and a published tiger haplotype assembly to identify structural rearrangements.
Figure 1 from Dean et al. 2026, BMC Genomics (CC BY 4.0). Telomere repeat positions and gaps across Sumatran tiger genome assemblies, showing 17 of 19 chromosomes achieving T2T level with ONT-only simplex read sequencing.
The hifiasm ONT assembly produced the highest contiguity, with error correction substantially improving both N50 and the proportion of the genome assembled into chromosome-scale contigs. After reference-guided scaffolding, 17 of the 19 tiger chromosomes achieved T2T-level assembly — compared to only 1 chromosome in the previously published tiger haplotype assembly. Two large chromosomal inversions were identified between the domestic cat and tiger genomes on chromosomes D4 and E2, involving genes related to glycoprotein degradation, immune function, and apoptosis — potential contributors to species-specific adaptations. Additionally, several structural rearrangements were discovered between the Sumatran tiger assembly and the existing tiger haplotype, particularly a large rearrangement on chromosome E1. De novo annotation predicted 23,737 complete genes, of which 20,511 matched Swiss-Prot entries.
This study demonstrates that a single ONT sequencing run — without PacBio HiFi, Illumina short reads, or additional scaffolding technologies — can yield near-T2T genome assemblies for mammalian genomes. For conservation genomics, this represents a significant advance: high-quality reference genomes can now be produced for endangered species at lower cost and with simpler workflows. The Sumatran tiger genome provides a foundational resource for population monitoring, inbreeding assessment, and captive breeding management. This approach directly informs our long-read sequencing service design for animal research projects — showing that ONT-only strategies are technically viable for species where budget or sample availability precludes multi-platform approaches.
The following representative data visualizations illustrate the assembly quality our clients can expect from long-read sequencing projects for animal and plant genomes. These demos are based on actual project outputs across diverse species.
Genome Assembly Contiguity Comparison. The N50 bar chart compares assembly contiguity achieved by short-read (Illumina), PacBio HiFi, and ONT ultra-long approaches for representative plant genomes (Arabidopsis thaliana ~135 Mb, Oryza sativa ~430 Mb, Sorghum bicolor ~730 Mb). While short-read assemblies plateau at N50 values of 20–50 kb, PacBio HiFi assemblies achieve 10–40 Mb N50, and ONT ultra-long assemblies reach 30–70 Mb depending on genome complexity.
BUSCO Completeness Assessment. BUSCO scores against the embryophyta_odb10 and metazoa_odb10 lineage databases consistently exceed 95% for long-read assemblies, compared to 85–92% for short-read assemblies of the same species. The improvement is most pronounced in the "fragmented" and "missing" categories — long reads resolve gene models that span repeat-rich intergenic regions where short reads fail to map.
Hi-C Contact Map. The Hi-C interaction heatmap demonstrates chromosome-scale scaffolding, with clear diagonal interaction blocks corresponding to individual chromosomes and minimal off-diagonal signal indicating low mis-join rates. T2T assemblies show interaction signal extending to chromosome ends, confirming telomere-to-telomere completeness.
Representative demo data: genome assembly N50 comparison across platforms, BUSCO completeness assessment, and Hi-C contact map demonstrating chromosome-scale scaffolding.
It depends on your genome characteristics and research goals. PacBio HiFi (Revio) delivers >Q30 accuracy with 15–25 kb reads, making it ideal for phased diploid and polyploid assembly of genomes under 3 Gb. Oxford Nanopore (PromethION) produces ultra-long reads (100+ kb, some exceeding 1 Mb) at lower per-base accuracy (~Q20), making it better for spanning large repeat arrays in giant genomes (>3 Gb), resolving centromeres and telomeres, and detecting DNA modifications natively. Many animal and plant genome projects achieve optimal results with both: PacBio for base-level phasing accuracy and ONT for ultra-long scaffolding. Our team provides a platform recommendation based on your species' genome size, ploidy, and repeat characteristics.
Yes. Polyploid genomes are a core area of our expertise. For allohexaploid wheat (~15 Gb), we deploy high-coverage PacBio HiFi sequencing for subgenome-resolved phasing, ONT ultra-long reads for gap closure, and Hi-C for chromosome-scale scaffolding. For autopolyploids like sugarcane, we use allele-aware assemblers and haplotype-resolved pipelines that account for multiple homologous copies. The key technical requirement is sufficient sequencing depth to distinguish subgenome-specific variants — we calculate coverage requirements based on genome size and ploidy during project design.
Short-read assemblies of complex genomes typically produce tens of thousands of contigs with N50 in the kilobase range, leaving unresolved gaps at repeat arrays, centromeric satellites, and segmental duplications. These gaps hide biologically important sequences — resistance gene clusters in plants, immune gene complexes in vertebrates, and structural variants underlying adaptive traits. Long-read assemblies routinely achieve megabase-scale contig N50, with T2T assemblies closing all gaps. The sorghum T2T assembly (Li et al., 2024) illustrates the difference: it eliminated 3,913 gaps from the previous reference, resolved all centromeres, and captured previously inaccessible pericentromeric gene families.
Yes — de novo sequencing of non-model species is our standard workflow. We perform whole-genome sequencing with the platform best suited to your species, followed by assembly with species-appropriate tools (hifiasm, Flye, NextDenovo). Without a reference, we estimate genome size and heterozygosity from k-mer frequency analysis to determine optimal sequencing coverage. The assembled genome is scaffolded with Hi-C data, quality-assessed with BUSCO against the appropriate lineage database, and annotated using evidence-guided gene prediction. For the most challenging cases — organisms with no close relative with an annotated genome — we build custom repeat libraries and train gene predictors from RNA-seq data generated in parallel.
Turnaround varies by genome size and project scope. Typical timelines: sample QC and HMW DNA extraction (3–5 business days), library preparation (3–5 business days), sequencing (5–14 days depending on platform and coverage), and bioinformatics (10–20 business days for standard assembly and annotation; longer for polyploid genomes or comparative multi-species analyses). Expedited service is available for time-sensitive projects. We provide a detailed timeline estimate during project consultation based on your specific species and requirements.
Yes. Oxford Nanopore sequencing directly detects 5-methylcytosine (5mC) and N6-methyladenine (6mA) from native DNA during sequencing — no bisulfite conversion or enzymatic treatment needed. The modification-specific current signal perturbations are identified computationally, producing genome-wide methylation maps at single-molecule resolution. PacBio HiFi reads also detect 5mC through analysis of inter-pulse duration in the polymerase kinetics signature, though with different sensitivity. For plant epigenomics, where DNA methylation regulates transposon silencing, imprinting, and stress-responsive gene expression, ONT direct methylation detection is particularly powerful because it preserves native DNA and provides simultaneous sequence + modification information from a single run.
Yes. Long-read sequencing is exceptionally well-suited for organelle genomes. PacBio HiFi circular consensus sequencing routinely assembles complete mitochondrial and chloroplast genomes in single contigs — resolving the structural complexity (repeat-mediated recombination, heteroplasmy, nuclear-mitochondrial DNA transfers) that fragments short-read organelle assemblies. We offer dedicated Mitochondrial/Chloroplast Genome Sequencing as part of our animal and plant genomics portfolio, with options for organelle-only projects or organelle assembly as an add-on to whole-genome sequencing.
To prepare an accurate project proposal, we typically need: species name and estimated genome size (or closest relative's genome size if unknown), ploidy level, research objectives (de novo assembly, resequencing, pan-genome, organelle, RNA-seq), sample type and estimated quantity available, and desired deliverables (assembly only, assembly + annotation, comparative genomics). Use the inquiry form below or contact our team directly — we respond within one business day with a preliminary proposal including platform recommendation, coverage strategy, timeline estimate, and deliverables scope.
References
For Research Use Only. Not for use in diagnostic procedures.