Complete genome assembly by PacBio HiFi and Oxford Nanopore PromethION — de novo assembly, reference-guided assembly, and hybrid strategies for microbial, fungal, plant, animal, and human genomes — producing chromosome-level, telomere-to-telomere assemblies with comprehensive structural and functional annotation
CD Genomics provides genome assembly services on PacBio HiFi (Revio/Sequel II), Oxford Nanopore PromethION, and Illumina NovaSeq platforms — from bacterial single-contig chromosomes and fungal telomere-to-telomere assemblies to plant and animal genomes with Hi-C chromosome-scale scaffolding, comprehensive structural and functional annotation, and NCBI submission support.
Genome assembly is the computational process of reconstructing complete genome sequences from the short DNA fragments (reads) produced by sequencing instruments. From the first bacterial genomes assembled by Sanger sequencing in the 1990s to the telomere-to-telomere (T2T) human genome completed in 2022, genome assembly has been transformed by advances in sequencing technology — most profoundly by the emergence of long-read sequencing platforms that produce reads tens to hundreds of kilobases in length, capable of spanning repetitive regions and structural complexities that were intractable to short-read methods.
At CD Genomics, we provide comprehensive genome assembly services across all major long-read and short-read platforms — PacBio HiFi (Sequel II / Revio), Oxford Nanopore (PromethION R10.4.1), Illumina NovaSeq, and Hi-C proximity ligation — supporting the full range of genome assembly projects: from bacterial and fungal genomes (1–100 Mb) to plant, animal, and human genomes (100 Mb to 20+ Gb). Our platform-agnostic approach ensures that each project receives the optimal sequencing and assembly strategy for its specific genome size, complexity, ploidy, and research objectives.
Whether you need a complete bacterial genome in a single contig, a haplotype-resolved assembly of a highly heterozygous plant, or a telomere-to-telomere reference genome for a non-model organism, our team of experienced bioinformaticians and molecular biologists will design and execute a tailored genome assembly solution — from DNA extraction and sequencing library preparation through assembly, curation, and comprehensive annotation — delivering publication-ready genome assemblies with full documentation.
Genome assembly — the process of reconstructing a complete genome sequence from sequencing reads — is the foundational step in nearly all genomics research. The quality of the genome assembly determines virtually everything that follows: the accuracy of gene prediction, the completeness of repeat annotation, the reliability of comparative genomics, and the interpretability of population genetic analyses. A fragmented or error-prone assembly propagates uncertainty through every downstream analysis; a complete, accurate assembly enables confident biological discovery.
The transition from short-read to long-read sequencing has been the single most transformative advance in genome assembly. Short-read platforms (Illumina) produce reads of 150–300 bp with very high accuracy (Q30+), but the fundamental limitation of read length means that repetitive regions longer than the read length cannot be uniquely resolved. Since most eukaryotic genomes contain 30–80% repetitive sequence — including transposable elements, tandem repeats, rDNA arrays, and centromeric satellites — short-read assemblies are necessarily fragmented into thousands or tens of thousands of contigs, with collapsed repeats, missing telomeres, and misassembled structural variants.
Long-read sequencing overcomes this limitation fundamentally. PacBio HiFi reads (15–25 kb, Q30+) provide the accuracy of short reads with 100× the read length, enabling single-copy resolution of most repetitive elements. Oxford Nanopore ultra-long reads (50–100+ kb, and up to 2 Mb) can span entire centromeric repeat arrays, large rDNA clusters, and complex structural rearrangements in a single contiguous read. When combined — HiFi for base-level accuracy and haplotype phasing, ultra-long ONT reads for bridging the largest repeats, and Hi-C for chromosome-scale scaffolding — modern long-read assembly pipelines routinely produce telomere-to-telomere genomes with zero gaps, fully resolved centromeres, and complete haplotype information.
At CD Genomics, we offer genome assembly services for organisms across the tree of life. Our service is organized by genome size and complexity tier — from microbial genomes (which we can complete in a single contig per chromosome) to large plant and animal genomes (where we integrate HiFi, ultra-long ONT, and Hi-C for chromosome-scale T2T assemblies) — with dedicated pipelines and bioinformatic expertise for each category.
Our integrated PacBio HiFi + ONT ultra-long + Hi-C pipeline routinely produces assemblies with zero spanned gaps, complete telomere resolution at every chromosome end, fully assembled centromeric repeat arrays, and complete rDNA clusters — not just the “best possible” assembly but the definitive genome sequence of the organism.
For heterozygous, diploid, and polyploid genomes, we offer a choice between phased (haplotype-resolved) and collapsed consensus assemblies. PacBio HiFi with hifiasm produces high-quality phased assemblies that separate haplotypes and resolve allelic differences — essential for studies requiring allele-specific expression, heterozygosity analysis, or accurate variant discovery.
PacBio HiFi and ONT sequencing natively detect DNA base modifications (5mC, 6mA, 4mC) from the same data used for genome assembly — providing simultaneous genome and epigenome characterization from a single sequencing run without additional experimental costs.
From 1 Mb bacterial genomes to 20+ Gb plant genomes, our sequencing capacity and bioinformatics infrastructure scale linearly. Multiplexed barcoding, flexible coverage targets, and tiered analysis packages ensure cost efficiency across projects of any scale.
We operate all major sequencing platforms in-house and recommend the optimal strategy based purely on the biology of your genome and research questions — not platform availability. This independence ensures you receive the best possible assembly for your investment.
Every project includes NCBI GenBank submission support, detailed methods sections suitable for genome announcement or resource papers, and publication-ready figures. Our team has extensive experience with genome assembly publications across diverse taxonomic groups.
Beyond reference-grade assemblies for individual organisms, long-read genome assembly is increasingly applied in pangenome construction and comparative genomics across populations, where complete, haplotype-resolved assemblies from multiple individuals of the same species enable discovery of gene presence-absence variation, structural variant landscapes, and accessory genome content that single-reference approaches systematically miss. International initiatives such as the Vertebrate Genomes Project (VGP), Earth BioGenome Project (EBP), and species-specific pangenome consortia (human, plant, livestock) increasingly rely on long-read assembly as the foundational technology for producing the reference-quality genomes that power cross-species comparative analysis and population genomics.
We provide dedicated genome assembly pipelines optimized for each major taxonomic group. Click on each service for detailed information.
Complete bacterial genome assembly from single-contig circular chromosomes to multi-replicon genomes with plasmid reconstruction. Our bacterial WGS de novo sequencing service delivers complete, circularized genomes with automatic plasmid detection and comprehensive functional annotation (COG, KEGG, GO, CARD, VFDB, CAZy, antiSMASH). For strains with an existing reference, our bacterial WGS resequencing service provides cost-effective variant discovery.
Telomere-to-telomere genome assembly for yeasts, filamentous fungi, and mushrooms — from compact 10 Mb yeast genomes to large repeat-rich filamentous fungal genomes exceeding 100 Mb. Our fungal WGS de novo sequencing service features fungal-specific annotation including CAZy carbohydrate-active enzymes, antiSMASH secondary metabolite clusters, and PHI-base effector prediction.
Chromosome-scale genome assembly for plants and animals, integrating PacBio HiFi, ultra-long ONT, and Hi-C proximity ligation for T2T pseudomolecules. Our services cover genome size estimation, repeat characterization, genome assembly, and comprehensive annotation — from small invertebrate genomes to large vertebrate and plant genomes exceeding 10 Gb. Contact us to discuss your project.
Complete genome sequencing and assembly for DNA and RNA viruses, combining short-read and long-read platforms for maximum coverage of viral genomes, quasispecies detection, and identification of recombinant or novel viral strains. Genome assembly from direct RNA sequencing available for RNA viruses without reverse transcription bias.
Each genome assembly project begins with a consultation to determine genome size, complexity, ploidy, and research objectives. High-molecular-weight DNA is extracted using taxon-specific protocols — enzymatic lysis for bacteria and fungi, CTAB-based for plants, proteinase K/SDS for animal tissues — with DNA quality and fragment length verified by pulsed-field gel electrophoresis, Qubit fluorometry, and TapeStation analysis.
Platform-specific libraries are prepared for the chosen sequencing strategy. PacBio HiFi: 15–20 kb SMRTbell libraries with HiFi CCS mode. ONT: Native ligation libraries (SQK-LSK114) for ultra-long reads, with optional ultra-long enrichment. Illumina: 350–550 bp paired-end libraries for genome survey or polishing. Hi-C: Proximity ligation libraries for chromosome-scale scaffolding. All libraries are barcoded for multiplexing where appropriate.
Sequencing is performed to target coverage optimized for the genome size and complexity. Coverage recommendations range from 30–100× for PacBio HiFi and 40–100× for ONT, scaled to genome size. Hi-C sequencing targets 100–150× for effective scaffolding.
Figure 1. Complete genome assembly workflow — from HMW DNA extraction and multi-platform sequencing through assembly (Flye, hifiasm, CANU), Hi-C scaffolding, polishing, T2T gap closure, and comprehensive genome annotation.
Assembly is performed using platform-appropriate algorithms: hifiasm (PacBio HiFi, with diploid phasing), Flye (ONT and HiFi), CANU / HiCanu (ONT ultra-long). Hi-C reads are used to scaffold contigs into chromosome-scale pseudomolecules using 3D-DNA / Juicebox or YaHS. Gap closure is performed using targeted ONT ultra-long reads or TGS-GapCloser. Manual curation with IGV and pretext contact maps ensures assembly correctness. Telomere repeat identification and centromere annotation complete the T2T validation. Assembly quality is assessed with QUAST, BUSCO, Merqury (k-mer completeness), and CheckM (for microbial genomes).
Structural annotation combines ab initio gene prediction, homology-based methods, and transcript evidence. Functional annotation uses kingdom-specific databases. Repeat annotation includes transposable element identification and classification. Deliverables include annotated GenBank files, complete genome maps, and a comprehensive project report.
Our bioinformatics pipeline supports comprehensive genome analysis from raw read processing through assembly, quality assessment, and multi-level annotation — organized across two service tiers.
| Analysis Feature | Basic Package | Advanced Package |
| Read QC & preprocessing | ✓ FastQC, NanoPlot, MultiQC; adapter trimming, quality filtering, length selection | ✓ + Contamination screening, k-mer analysis (GenomeScope), genome size estimation |
| De novo assembly | ✓ Flye / hifiasm / CANU (platform-optimized); initial contig assembly | ✓ + Multi-assembler comparison; Hi-C scaffolding (3D-DNA / YaHS); haplotype-phased assembly (hifiasm) |
| Assembly polishing & gap closure | ✓ Pilon (Illumina-based), Medaka (ONT-based), or Polypolish | ✓ + TGS-GapCloser targeted gap closure; manual IGV curation; telomere & centromere identification |
| Assembly quality assessment | ✓ QUAST (N50, L50, contig count, total length), BUSCO (lineage-specific) | ✓ + Merqury k-mer completeness; CheckM (microbial); contact map validation (Hi-C); CRAQ quality score |
| Repeat annotation | ✓ RepeatModeler + RepeatMasker; TE class/family classification | ✓ + TE insertion time estimation; landscape analysis; proximity to genes; RIP index (fungi) |
| Structural gene annotation | ✓ BRAKER3 / Augustus / GeneMark-ES with transcript evidence; tRNA/rRNA prediction | ✓ + Manual curation of problematic loci; alternative isoform detection; UTR annotation; ncRNA identification |
| Functional annotation (general) | ✓ InterProScan, Pfam, GO, KEGG, COG, Nr alignment | ✓ + KEGG pathway mapping; Enzyme Commission numbers; Transporter classification (TCDB) |
| Kingdom-specific annotation | ✓ Standard databases (RefSeq, UniProt) | ✓ CAZy, antiSMASH, PHI-base, DFVF (fungi); PLAZA/Ensembl Plants (plants); VEGA/RefSeq (animals); CARD/VFDB (microbes) |
| Comparative genomics | — | ✓ Orthofinder gene family analysis; phylogenomics; synteny visualization (MCscanX); pan-genome construction |
| Epigenome analysis | — | ✓ 5mC/6mA/4mC detection from HiFi or ONT data; methylation motif analysis; methylome comparison |
| Custom reporting & visualization | ✓ Standard project report PDF with assembly statistics and annotation tables | ✓ Interactive genome browser (JBrowse2); Circos genome maps; publication-ready figures; NCBI submission files |
The optimal strategy for a genome assembly project depends on genome size, complexity, ploidy, and research objectives. Our team provides platform-neutral recommendations tailored to each project.
| Strategy | PacBio HiFi-Only | ONT Ultra-Long-Only | HiFi + ONT + Hi-C (Integrated) |
| Best for genome types | Microbial, fungal, small plant/animal (<500 Mb) | Large repeat-rich genomes; microbial; rapid screening | Large plant/animal (>500 Mb); T2T reference-grade; complex polyploid genomes |
| Assembly contiguity | ★★★★☆ (T2T for small genomes) | ★★★★★ (reads span largest repeats) | ★★★★★ (multi-platform synergy) |
| Base accuracy | ★★★★★ (Q30+ HiFi) | ★★★☆☆ (Q14–Q20 raw; Q30+ polished) | ★★★★★ (HiFi-polished) |
| Haplotype phasing | ★★★★★ (hifiasm phasing) | ★★★☆☆ (limited) | ★★★★★ (HiFi phasing + ONT contiguity) |
| T2T completion | ★★★☆☆ (limited by read length for large genomes) | ★★★★☆ (ultra-long spans repeats) | ★★★★★ (ONT gaps → HiFi polishing → Hi-C scaffolding) |
| Hi-C scaffolding needed? | Usually for >100 Mb genomes | Usually for >500 Mb genomes | Yes (chromosome-scale pseudomolecules) |
| Coverage (small genome <100 Mb) | 30–50× CCS | 40–80× raw | 30× HiFi + 40× ONT |
| Coverage (large genome >1 Gb) | 30–60× CCS | 30–60× raw | 30× HiFi + 20× ONT + 100× Hi-C |
| Per-genome cost | $$–$$$ (moderate to high) | $ (lowest) | $$$–$$$$ (highest, best quality) |
For most reference-genome projects, we recommend the integrated HiFi + ONT + Hi-C strategy as the gold standard. For smaller genomes or budget-constrained projects, single-platform strategies (HiFi-only for accuracy, ONT-only for cost efficiency) produce excellent results. Contact our team for a free project consultation and strategy recommendation.
| Category | Requirement | Notes |
| Sample type | Genomic DNA (fresh or frozen tissue, cell pellet, blood, microbial culture, plant leaf) | DNA extraction service available for challenging samples (polysaccharide-rich, low-biomass, or degraded) |
| Minimum input (gDNA) | 500 ng (Illumina); 5–10 µg (PacBio HiFi); 5–10 µg (Nanopore) | HMW DNA (≥ 30 kb fragments) strongly recommended for long-read assembly; minimum input depends on genome size and library type |
| DNA quality | OD260/280: 1.8–2.0; OD260/230: ≥ 1.8; no visible degradation; HMW verified by PFGE | Taxon-specific DNA purification methods used to remove species-specific contaminants (polysaccharides, polyphenols, humic acids) |
| Genome information | Estimated genome size, ploidy, G+C content, and heterozygosity (if known) | Pilot sequencing (low-coverage Illumina or ONT) available for genome size and complexity estimation for novel organisms |
| Sample numbers | Single individual to population-scale studies | Multiplexed barcoding enables cost-effective assembly of multiple samples; large projects receive dedicated project management |
| Shipping conditions | gDNA: ice pack (4°C) or dry ice; Tissue/pellet: dry ice or liquid nitrogen | See our Sample Submission Guidelines for detailed instructions |
Independent Platform Expertise
We operate all major sequencing platforms in-house — PacBio HiFi (Sequel II / Revio), Oxford Nanopore (PromethION R10.4.1), Illumina (NovaSeq), and Hi-C — and our recommendations are based purely on the scientific requirements of your project. This independence is especially important for genome assembly, where the optimal strategy varies dramatically with genome size, repeat content, ploidy, and budget.
Proven Track Record Across Kingdoms
Our team has delivered genome assembly projects across the full taxonomic spectrum: from microbial genomes (single-contig circular chromosomes) to fungal genomes (with phased diploid assemblies and secondary metabolite annotation) to plant and animal genomes (with Hi-C chromosome-scale scaffolding and T2T completion).
T2T & Haplotype-Resolved Assembly Expertise
Our bioinformatics team maintains deep expertise in the latest assembly algorithms — hifiasm for diploid phasing, Flye for ONT and HiFi assembly, Hi-C scaffolding pipelines, and manual curation workflows — and actively tracks developments in the rapidly evolving genome assembly field to ensure our clients benefit from the best available methods.
Comprehensive Post-Assembly Support
A genome assembly project does not end with a FASTA file. We provide full downstream support including structural and functional annotation, comparative genomics, NCBI submission, and methods writing for publications.
Wang F, Bao J, Zhang H, Zhai G, Song T, Liu Z, Han Y, Yu F, Zou G, Zhu Y. A telomere-to-telomere genome assembly of Chinese grain sorghum 654. Scientific Data. 2025;12:460. doi:10.1038/s41597-025-04791-6.
Sorghum (Sorghum bicolor) is the world’s fifth most important cereal crop and a key model for C4 photosynthesis research, drought tolerance, and bioenergy production. The existing reference genome for sorghum (BTx623) was a landmark assembly but contained approximately 4,000 spanned gaps — unresolved regions including complex centromeric repeats, rDNA arrays, and telomeric sequences — that limited comprehensive genomic analysis. The Chinese grain sorghum inbred line 654 is an elite variety with important agronomic traits, but no complete genome sequence was available for this or any sorghum line.
In this study, Wang et al. generated a complete, gap-free T2T genome assembly of sorghum line 654 using an integrated multi-platform long-read sequencing strategy — demonstrating the current state of the art in plant genome assembly.
The genome was assembled using three complementary sequencing technologies: PacBio HiFi (Revio, 36.4× coverage, 15–20 kb reads), ONT ultra-long reads (PromethION R10.4.1, 24.3× coverage, N50 read length >50 kb), and Hi-C proximity ligation (95.7× coverage). Initial assembly was performed with hifiasm using HiFi reads, scaffolded to chromosome-level with Hi-C data, and gaps filled using ONT ultra-long reads. Assembly quality was validated by multiple complementary methods including k-mer completeness (Merqury), BUSCO gene completeness, Hi-C contact map consistency, and manual inspection of all telomere and centromere junctions.
Figure 2. Telomere-to-telomere genome assembly of sorghum line 654 using integrated PacBio HiFi + ONT ultra-long + Hi-C strategy. The assembly (728.81 Mb) achieved zero gaps, all 10 centromeres and 20 telomeres fully resolved, with a QV score of 64.72 and 44,399 annotated protein-coding genes. Adapted from Wang et al. (2025), Scientific Data, CC BY 4.0.
This study demonstrates that integrated multi-platform long-read sequencing — combining the accuracy of PacBio HiFi with the repeat-spanning capability of ONT ultra-long reads and the chromosome-scale information of Hi-C — can produce complete, telomere-to-telomere plant genome assemblies of the highest quality. The approach and coverage specifications serve as a benchmark for plant genome assembly projects: ~36× HiFi for accuracy, ~24× ONT ultra-long for gap closure, and ~100× Hi-C for scaffolding. The methods are directly applicable to other plant and animal genomes of similar or larger size.
CD Genomics provides free project consultation to help determine the optimal genome assembly strategy for your specific organism and research objectives. Contact our scientists to discuss your project requirements.
De novo genome assembly constructs a genome sequence from sequencing reads alone, without using any existing reference genome as a guide. It produces the complete genome sequence of the organism and is required for species without a high-quality reference. Reference-guided assembly (resequencing) aligns reads to an existing reference genome to identify differences such as SNPs, InDels, and structural variants. De novo assembly is the appropriate choice when studying a species without a reference genome, characterizing novel structural elements, or building a complete genome blueprint. Reference-guided assembly is more cost-effective when a high-quality reference is available and the goal is variant discovery rather than genome reconstruction.
Assembly completeness depends on genome size, complexity, and the chosen sequencing strategy. For microbial genomes (bacteria, archaea), we routinely achieve single-contig complete chromosomes. For fungal genomes (10–100 Mb), telomere-to-telomere assemblies with zero gaps are standard with our integrated HiFi + ONT pipeline. For small to moderate plant and animal genomes (100–500 Mb), chromosome-scale assemblies with <50 gaps are typical. For large and complex genomes (>1 Gb, high repeat content), assemblies typically achieve chromosome-scale pseudomolecules with 50–200 remaining gaps. The integrated HiFi + ONT + Hi-C strategy delivers the highest completeness across all genome types. Specific targets are defined during project design.
Yes. PacBio HiFi reads with hifiasm produce high-quality phased assemblies for diploid and polyploid genomes. For highly heterozygous species, we offer two strategies: (1) haplotype-resolved assembly that separates the two haplotypes into distinct phased pseudomolecules, and (2) collapsed haploid assembly that produces a single consensus sequence with heterozygous sites represented in the ambiguity code. The choice depends on your research objectives. Hi-C data is particularly valuable for polyploid genome assembly, as it helps assign phased contigs to the correct subgenome.
We report a comprehensive suite of assembly quality metrics including: (1) contiguity metrics — total assembly size, number of contigs/scaffolds, N50, L50, longest contig, number of gaps; (2) completeness metrics — BUSCO completeness (lineage-specific odb10 database), Merqury k-mer completeness, CheckM completeness (for microbial genomes); (3) accuracy metrics — QV score (consensus accuracy), number of structural errors identified by Merqury; (4) chromosome-scale metrics — number of telomeres identified, number of centromeres resolved, Hi-C contact map quality. These metrics collectively provide a rigorous and transparent assessment of assembly quality, comparable to published reference genome standards.
For de novo assembly projects, no reference genome is needed — the goal is to create one. For reference-guided assembly (resequencing) projects, we can work with a client-provided reference or help identify the most suitable publicly available reference genome from NCBI RefSeq, Ensembl, or other databases. If no suitable reference exists for the target species, we can first generate a high-quality de novo assembly (which then serves as the reference for resequencing). Our project consultation includes guidance on the best approach for your specific research question and organism.
Deliverable Examples for Genome Assembly Projects
1. Complete genome assembly file (FASTA) with chromosome-level pseudomolecules — one sequence per chromosome with telomeric repeats at both ends, plus mitochondrial genome and any accessory or extrachromosomal elements.
2. Genome annotation file (GFF3/GenBank format) with structural and functional annotations — protein-coding genes, tRNA, rRNA, ncRNA, repeat annotations, and predicted functional domains with cross-references to kingdom-specific databases.
3. Hi-C contact map and chromosome-scale scaffolding validation — interactive heatmap showing chromatin interaction patterns and evidence supporting chromosomal pseudomolecule assignments.
4. Comprehensive project report PDF documenting experimental methods, sequencing QC, assembly statistics, annotation summary, and detailed methods text suitable for manuscript preparation.
Figure 3. Representative deliverable formats for genome assembly projects. Left: chromosome-scale genome map with multi-track annotation. Center: Hi-C contact map validating chromosome-scale scaffolding. Right: assembly quality metrics and functional annotation summary. AI-generated representative data.
References
For research use only. Not for use in diagnostic procedures.