CD Genomics provides long-read sequencing-based comparative genomic analysis on PacBio Revio and ONT PromethION platforms, delivering complete, phased genome assemblies that capture the full spectrum of genomic variation across species. Our end-to-end service covers everything from HMW DNA extraction through multi-genome comparative bioinformatics — including whole-genome alignment, structural variant discovery, transposable element annotation, and phylogenomic reconstruction.
Most comparative genomic studies today still rely on next-generation sequencing (NGS) short reads that fragment at repetitive elements, collapse segmental duplications, and miss the large structural variants that drive species differentiation. CD Genomics takes a fundamentally different approach. We specialize exclusively in long-read sequencing (third-generation sequencing) for comparative genomic analysis — using PacBio Revio HiFi and ONT PromethION platforms to deliver complete, phased genome assemblies that capture the full spectrum of genomic variation across species.
Our long-read comparative genomics service is designed for researchers who need more than gene-by-gene comparisons. We provide end-to-end solutions: from high-molecular-weight DNA extraction and long-read library preparation through PacBio Revio or ONT PromethION sequencing to comprehensive bioinformatics analysis including whole-genome alignment, structural variant discovery, transposable element annotation, and phylogenomic reconstruction. This is not a hybrid approach that defaults to short reads — third-generation sequencing is our core technology platform, and that makes all the difference for comparative genomics.
Comparative genomic analysis examines the complete DNA sequences of multiple species, populations, or individuals to identify similarities, differences, and evolutionary relationships. The resolution and accuracy of these comparisons depend fundamentally on the quality of the underlying genome assemblies — and this is precisely where the choice of sequencing technology makes or breaks a comparative genomics project.
The gap between short-read and long-read comparative genomics is not incremental — it is a difference in what biological questions can be asked and answered. Below is a direct comparison of key metrics that define the resolution of comparative genomic analysis on each platform.
| Comparative Genomics Metric | NGS Short-Read (150–300 bp) | Long-Read (15–100+ kb) |
| Typical assembly contiguity (N50) | < 1 Mb; often < 100 kb for non-model species | > 50 Mb; routinely chromosome-scale |
| Repetitive region representation | Collapsed or absent — repeats > read length are unresolvable | Fully spanned and resolved, including full-length TEs and segmental duplications |
| Structural variant detection (cross-species) | 15–30% sensitivity — misses inversions, large indels, complex rearrangements | > 90% sensitivity — detects SVs from 50 bp to megabase scale |
| Haplotype phasing | Statistical inference from population data — unreliable for single individuals or non-model species | Direct single-molecule phasing — resolves haplotypes for any individual without population data |
| Transposable element annotation | Fragmentary — only consensus TE sequences recoverable | Full-length TE annotation — insertion ages, copy numbers, mobilization histories |
| Genome completeness (BUSCO) | 70–85% — gene sets incomplete, regulatory regions often missing | 90–98% — near-complete gene and regulatory landscape |
| Orthologous gene family resolution | Gene copies collapsed — copy-number variation underestimated | Gene family expansions and contractions fully resolved |
These differences translate directly into biological conclusions. A comparative genomics study built on NGS assemblies may conclude that two species share similar gene family content — when in reality, lineage-specific expansions in repetitive genomic regions were simply collapsed in the assembly and never detected. Similarly, inferences about genome size evolution, transposable element dynamics, and structural rearrangements are fundamentally unreliable when drawn from fragmented short-read assemblies. Third-generation sequencing eliminates this systematic bias, providing a complete and unbiased foundation for comparative inference.
Our service operationalizes this advantage through three integrated capabilities. De novo genome assembly for multiple species — PacBio Revio HiFi reads provide highly accurate (≥Q30) consensus sequences for gene annotation and variant discovery, while ONT PromethION ultra-long reads bridge complex repeats for maximum contiguity. Our dedicated genome assembly services deliver complete, phased assemblies for each species in your comparative study. Cross-species comparative analysis — we align whole genomes to identify orthologous regions, detect conserved synteny blocks, characterize lineage-specific indels, and reconstruct phylogenetic relationships at the genome level. Population-scale genomic comparison — for studies involving multiple individuals per species, our long-read approach resolves population-specific structural variants and haplotypes that short-read population genomics routinely misses.
Comparative genomics powered by third-generation sequencing is the systematic comparison of complete, contiguous genome sequences across two or more species, strains, or individuals to understand genome evolution, functional divergence, and the genetic basis of phenotypic variation. Unlike approaches that rely on mapping short reads to a single reference genome — which introduces reference bias and misses reference-absent variation — long-read comparative genomics builds high-quality de novo assemblies for each entity being compared, ensuring that every genome is represented on its own terms.
Long-read sequencing addresses the fundamental limitation that has constrained comparative genomics for two decades: the inability of short reads to resolve repetitive DNA. In a typical mammalian genome, more than half of the sequence consists of repetitive elements, and these repeats are precisely the regions that drive genome size evolution, facilitate gene duplication, and generate structural variation. When these regions are collapsed or missing in assemblies, comparative analyses of gene family evolution, transposable element dynamics, and non-coding regulatory sequences are fundamentally compromised. Long-read sequencing eliminates this limitation by providing reads that span entire repetitive elements, enabling complete genome assemblies that capture the full spectrum of genomic features that evolve between species.
Long-read assemblies capture repetitive regions, segmental duplications, and GC-rich sequences systematically under-represented in NGS assemblies. This completeness transforms downstream comparative analyses — gene family evolution, transposable element dynamics, and non-coding regulatory evolution cannot be accurately studied from fragmented assemblies.
Large insertions, deletions, inversions, and translocations are primary drivers of phenotypic divergence and speciation. Long-read sequencing detects structural variants from 50 bp to megabase scale with high sensitivity — short-read methods miss 70–85% of these events depending on size and genomic context.
Diploid and polyploid genomes contain multiple haplotypes that can differ substantially in gene content and regulatory architecture. Long reads phase haplotypes directly from single-molecule data, enabling allele-specific comparative analysis that computational phasing from short reads cannot recover.
We offer both PacBio Revio (highest single-base accuracy for gene-level analysis) and ONT PromethION (ultra-long reads for maximum contiguity and repeat resolution). For projects requiring both, our integrated hybrid strategy combines HiFi accuracy with ultra-long contiguity in a unified workflow.
Our analysis pipelines go beyond assembly to deliver whole-genome alignments, synteny maps, phylogenetic trees, gene family evolution analysis, and customized comparative genomics reports — not just raw sequence data.
We have extensive experience with simultaneous assembly and comparison of multiple genomes — from closely related strains or cultivars to phylogenetically diverse species across major eukaryotic lineages including animals, plants, fungi, and protists. Our pan-genome analysis workflows are designed for projects comparing multiple individuals or closely related species.
Beyond standard phylogenomic and repeat evolution analysis, long-read assemblies enable comparative epigenomic analysis using PacBio kinetic or ONT current-based methylation detection across species — allowing researchers to correlate epigenetic landscape divergence with genome evolution, a dimension of comparative genomics that remains inaccessible from short-read assemblies. This integrated genetic-epigenetic comparative framework is increasingly relevant for studies of phenotypic evolution, adaptation, and speciation.
The foundation of successful long-read comparative genomics is ultra-high-molecular-weight (UHMW) DNA. We extract HMW DNA from each species or individual using optimized protocols with fragment length analysis by FEMTO Pulse or equivalent. DNA quality is assessed by spectrophotometry (A260/280 ≥ 1.8) and fluorometric quantification, with a minimum fragment size of 30 kb recommended for PacBio Revio and 50+ kb for ONT PromethION ultra-long libraries. For non-model organisms, we optimize extraction protocols to overcome species-specific challenges such as polysaccharides, polyphenolics, or high nuclease content.
For PacBio Revio: SMRTbell library preparation using the SMRTbell Prep Kit 3.0. HMW DNA is sheared to 15–25 kb, end-repaired, and ligated to SMRTbell adapters. Libraries are size-selected using the BluePippin system to remove short fragments, then bound to polymerase and loaded onto SMRT Cell 8M for HiFi sequencing. Each Revio run produces approximately 90 Gb of HiFi reads (≥Q30), sufficient for 30× coverage of a 3 Gb genome in a single SMRT Cell.
For ONT PromethION: Ultra-long library preparation using the Ligation Sequencing Kit (SQK-LSK114) or Ultra-Long DNA Sequencing Kit. Sequencing is performed on R10.4.1 flow cells with the PromethION P48 or P24 compute module. Ultra-long reads routinely exceed 100 kb, with maximum read lengths surpassing 2 Mb. Each PromethION flow cell generates 100–290 Gb of data depending on library quality and run duration. For multi-species projects, native barcoding enables cost-effective multiplexed sequencing.
End-to-end long-read comparative genomics workflow from sample preparation to multi-genome comparative analysis.
We assemble each genome using platform-optimized assemblers: Hifiasm or HiCanu for PacBio HiFi reads, Flye or Shasta for ONT reads, or hybrid approaches combining both data types. Assemblies are polished using the same long-read data and evaluated for completeness (BUSCO), contiguity (N50, L50, number of contigs), and base accuracy (QV). For diploid and polyploid species, we perform phased assembly to resolve haplotype-specific differences, providing separate haplotype sequences for allele-aware comparative analysis.
With high-quality assemblies complete, our bioinformatics team conducts comprehensive comparative analysis: whole-genome alignment using minimap2 and MUMmer, synteny block detection and visualization, orthologous gene family clustering with OrthoFinder, phylogenetic tree reconstruction, transposable element annotation using RepeatModeler and RepeatMasker, and lineage-specific genomic feature identification. Every analysis is customized to the biological questions driving each project, with results delivered in publication-ready format.
| Analysis Feature | Basic | Advanced |
| De novo genome assembly (per species) | ✓ Single assembler | ✓ Multi-assembler optimization |
| Assembly quality assessment (BUSCO, QV, N50) | ✓ | ✓ |
| Repeat annotation and TE characterization | ✓ ab initio | ✓ Custom TE library + comparative dynamics |
| Gene prediction and annotation | ✓ ab initio + homology | ✓ Evidence-based (RNA-seq / Iso-Seq integration) |
| Whole-genome multi-species alignment | — | ✓ Pairwise + synteny visualization |
| Orthologous gene family clustering | — | ✓ OrthoFinder + GO/functional enrichment |
| Phylogenetic reconstruction | — | ✓ Maximum likelihood / coalescent-based trees |
| Cross-species structural variant detection | — | ✓ SV discovery and functional impact annotation |
| Comparative genomics report | ✓ Standard summary | ✓ Customized with publication-ready figures |
| Custom downstream analysis | — | ✓ Tailored to project-specific biological questions |
The choice between PacBio Revio, ONT PromethION, or a combined approach depends on your specific comparative genomics goals, genome characteristics, and budget. Our team provides platform-agnostic guidance to help you select the optimal strategy for your project.
| Feature | PacBio Revio | ONT PromethION | Hybrid (Revio + PromethION) |
| Read length | 15–25 kb (HiFi CCS) | 20–100+ kb (ultra-long) | Both ranges available |
| Base accuracy (consensus) | ≥Q30 (99.9%) | Q20+ (simplex) / Q30+ (duplex) | HiFi accuracy + ultra-long contiguity |
| Throughput per run | ~90 Gb (SMRT Cell 8M) | 100–290 Gb (flow cell) | Maximum combined output |
| Best suited for | Accurate gene annotation, SNV/indel calling, isoform analysis | Maximum contiguity, spanning complex repeats, resolving large SVs | Reference-grade assemblies, comprehensive multi-species comparison |
| Methylation detection | ✓ 5mC from HiFi kinetics | ✓ 5mC, 5hmC, 6mA native detection | Both modification types |
| Multi-species barcoding | ✓ Yes | ✓ Yes (native barcoding) | Flexible per species |
| Category | Requirement | Notes |
| Sample type | High-quality genomic DNA (tissue, blood, cells); or fresh/frozen tissue for extraction | For non-model species, we recommend tissue over extracted DNA for optimal HMW yields |
| Minimum input (DNA) | 1–5 µg HMW DNA (1 µg for PacBio; 5 µg for ONT ultra-long) | Lower inputs possible for HiFi libraries with amplification |
| DNA quality | A260/280 ≥ 1.8; A260/230 ≥ 1.8; no visible degradation | Integrity assessed by FEMTO Pulse; fragments ≥ 30 kb recommended |
| Number of species | Flexible — single pair to dozens | Multi-species projects benefit from barcoded multiplexing |
| Shipping | Overnight on dry ice or ice packs | Contact for sample-specific shipping recommendations |
Long-read is our specialty, not a side service.
Our comparative genomics platform is built from the ground up around third-generation sequencing. We do not default to short reads, and we do not force your project into a hybrid NGS+long-read workflow that adds complexity without commensurate benefit. Every species in your comparative study is sequenced and assembled using the same long-read standard, ensuring consistent data quality and comparability across genomes.
Dual-platform expertise under one roof.
We operate both PacBio Revio and ONT PromethION systems in-house, with experienced teams optimizing library preparation, sequencing, and analysis for each platform. We help you choose the right platform — or the right combination — for your specific comparative genomics questions, without platform bias.
Comparative genomics-dedicated bioinformatics.
Our bioinformatics team specializes in multi-genome comparative analysis, not just single-genome assembly. From whole-genome alignment and synteny mapping to transposable element dynamics and phylogenomics, we deliver analysis that directly addresses the biological questions driving your comparative study.
Proven track record across diverse taxa.
We have delivered successful long-read genome assemblies for species spanning mammals, birds, fish, insects, plants, fungi, and microbial eukaryotes, providing the cross-lineage experience needed to anticipate and overcome species-specific challenges in DNA extraction, library preparation, and assembly.
Guiglielmoni N, Villegas LI, Kirangwa J, Schiffer PH. Revisiting genomes of non-model species with long reads yields new insights into their biology and evolution. Frontiers in Genetics. 2024;15:1308527.
Genome assemblies of non-model organisms built from short-read sequencing are often highly fragmented, with thousands of contigs that collapse repetitive regions and misrepresent genomic architecture. The authors selected two nematode species — a diploid insect parasite (Romanomermis culicivorax, Mermithidae) and a triploid free-living species (Panagrolaimus sp. PS1159) — that had previously been assembled from short reads, to assess what biological insights are gained by revisiting them with long-read sequencing for comparative analysis.
Both species were sequenced using PacBio HiFi and ONT R10.4.1 long-read platforms. Assemblies were generated with multiple assemblers (Hifiasm, Flye, CANU) and compared to the original short-read assemblies. Comparative analyses included gene prediction, transposable element annotation, and haplotype divergence assessment across the two species.
Figure 3 from Guiglielmoni et al. 2024 (CC BY 4.0). Comparison of assemblies based on TE count and BUSCO ortholog statistics shows higher repeat and gene completeness of long-read assemblies versus the original short-read assemblies.
Long-read assemblies dramatically improved contiguity: the Mermithidae assembly was reduced from thousands of short-read contigs to just tens of long-read contigs with N50 exceeding 1 Mb. The long-read assemblies revealed substantially more transposable element content — TE families that were collapsed or entirely missing in the short-read assemblies — and enabled accurate annotation of haplotype divergence in the diploid species. For the triploid Panagrolaimus sp. PS1159, the phased long-read assembly provided three distinct haplotypes, revealing copy-number variation across orthologs that was completely invisible in the collapsed short-read assembly.
The study demonstrates that long-read sequencing fundamentally changes the quality — and therefore the biological conclusions — of comparative genomic analyses in non-model organisms. Key biological features including transposable element dynamics, haplotype structure, and gene family content that are central to evolutionary inference were poorly represented or entirely absent from the short-read assemblies, highlighting the critical importance of third-generation sequencing for accurate comparative genomics.
CD Genomics provides free project consultation to help determine the optimal comparative genomics strategy for your specific research questions and species of interest. Contact our scientists to discuss your project requirements.
There is no fixed upper limit. We have supported projects ranging from pairwise genome comparisons to multi-species studies spanning dozens of genomes. Barcoded multiplexing for both PacBio Revio and ONT PromethION allows cost-effective processing of multiple species simultaneously. For projects involving ten or more species, we recommend a tiered strategy with HiFi sequencing for representative high-priority genomes and optimized coverage for broader taxonomic sampling.
Not necessarily. For projects focused on accurate gene annotation and coding-sequence evolution, PacBio Revio HiFi alone provides excellent results. For projects requiring maximum contiguity — resolving complex repeats, centromeres, or large structural variants — ONT PromethION ultra-long reads are advantageous. Our team provides platform-agnostic guidance based on your genomes' specific characteristics, project scale, and biological questions.
Yes — this is our primary use case. Long-read de novo assembly does not require an existing reference genome. We assemble each species from scratch using its own long-read sequencing data, then conduct reference-free comparative analysis through whole-genome alignment and orthology mapping. This approach is particularly valuable for non-model organisms where no close reference is available.
The standard package includes de novo assembly for each species, assembly quality metrics (BUSCO, QV, N50), repeat annotation, ab initio gene prediction, and a comparative genomics summary report. The advanced package adds whole-genome multi-species alignment with synteny visualization, orthologous gene family clustering with functional enrichment, phylogenetic reconstruction, cross-species structural variant detection, and comprehensive transposable element comparative dynamics.
1. De novo genome assembly statistics for each species — contig N50, total assembly size, BUSCO completeness, and QV score
2. Whole-genome alignment dot plots and synteny maps between compared species
3. Orthologous gene family clustering results with functional annotation enrichment analysis
4. Comparative transposable element landscape across species — TE class and order composition, copy number, and estimated insertion ages
References
For research use only. Not for use in diagnostic procedures.