High-quality complete bacterial genome assembly, circular chromosome closure, plasmid reconstruction, and comprehensive structural and functional annotation by PacBio HiFi CCS, Oxford Nanopore PromethION, and hybrid Illumina+long-read strategies — for pathogenic bacteria, environmental isolates, probiotics, and industrial strains of any G+C content or genomic complexity
CD Genomics provides bacterial whole-genome de novo sequencing on PacBio HiFi (Q30+ CCS), Oxford Nanopore PromethION, and Illumina platforms, delivering complete, closed circular genomes and plasmids for any bacterial species. Our tri-platform pipeline spans HMW DNA extraction through assembly, polishing, and comprehensive functional annotation against 12+ databases — producing publication-ready genomes suitable for GenBank submission without manual gap closure.
Bacterial whole-genome de novo sequencing constructs a complete genome sequence from raw sequencing data without requiring a pre-existing reference genome. Unlike resequencing — which maps reads to an existing reference and identifies differences — de novo assembly builds the genome from scratch, making it indispensable for characterizing novel bacterial isolates, completing gap-free assemblies of known species, and resolving complex genomic regions such as ribosomal operons, mobilome elements, and large structural variants that are invisible to short-read approaches.
At CD Genomics, we provide bacterial whole-genome de novo sequencing on all three major sequencing platforms — PacBio HiFi (Sequel II / Revio), Oxford Nanopore (PromethION R10.4.1), and Illumina NovaSeq — as well as hybrid strategies that combine the strengths of multiple platforms. For bacterial resequencing applications, see our Bacterial Whole-Genome Resequencing service. Our bioinformatics pipelines deliver assembled, annotated, and QC-validated bacterial genomes with circular chromosomes, complete plasmid sequences, and comprehensive functional annotation ready for publication and GenBank submission.
Bacterial whole-genome de novo sequencing is the process of determining the complete DNA sequence of a bacterial genome through computational assembly of raw sequencing reads into contiguous sequences (contigs), and ultimately into complete circular chromosomes and plasmids, without reliance on a reference genome. The quality of the final assembly is fundamentally determined by the read length and accuracy of the sequencing platform used — which is why long-read technologies have transformed bacterial genomics.
Short-read sequencing (Illumina), while cost-effective and accurate per base, produces reads of only 150–300 bp — far shorter than the repetitive elements, insertion sequences, ribosomal RNA operons, and phage integrations that populate even compact bacterial genomes. These features create assembly gaps that fragment the genome into dozens or hundreds of contigs, obscuring the true genomic architecture and preventing complete genome closure. Long-read sequencing platforms — PacBio HiFi (15–25 kb CCS reads at Q30+) and Oxford Nanopore (reads exceeding 100 kb) — span these repetitive regions in single continuous reads, enabling assemblers to resolve complex genomic structures and produce complete, single-contig circular chromosomes.
At CD Genomics, we have established bacterial whole-genome de novo sequencing workflows across all major platforms, with extensive experience spanning over 500 bacterial species — from well-characterized pathogens (Escherichia coli, Staphylococcus aureus, Mycobacterium tuberculosis, Salmonella enterica) to environmental isolates, extremophiles, probiotics, and industrially significant strains across the full taxonomic breadth of the bacterial domain.
Long reads spanning repetitive elements (rRNA operons, IS elements, phage integrations, homopolymer tracts) enable assemblers to produce single-contig circular chromosomes and closed plasmids without manual gap closure — a task that is frequently impossible with short-read data alone.
Transposable elements, genomic islands, prophage insertions, and tandem gene duplications — all major drivers of bacterial evolution and adaptation — are correctly assembled rather than collapsed or deleted, preserving the true gene content and organization of the genome.
Complete assemblies from multiple strains of the same species capture the full gene repertoire — including strain-specific accessory genes, antimicrobial resistance determinants, and virulence factors that may be missed in fragmented draft genomes.
We operate PacBio, ONT, and Illumina platforms in-house and provide unbiased recommendations — from pure long-read assemblies for the highest contiguity to cost-optimized hybrid strategies for large-scale population studies — matched to your specific quality and budget requirements.
Whether you need a single reference-quality genome for a novel isolate or high-throughput de novo assembly of 500+ strains for a population genomics study, our multiplexed barcoding strategies and multi-platform capacity scale to meet project demands.
For industrial and pharmaceutical projects, we provide full traceability documentation including sample handling records, sequencing QC metrics, assembly quality statistics, and annotated genome files compatible with regulatory submission requirements.
Bacterial de novo sequencing is increasingly applied in One Health AMR surveillance programs, where complete genome assemblies from pathogens, livestock, and environmental isolates are compared to track resistance gene transmission across the human-animal-environment interface — a use case that fragmented draft genomes cannot reliably support.
Our extraction and library protocols are tailored to bacterial cell wall type — enzymatic lysis for Gram-positive and high-G+C actinobacteria, alkaline lysis for Gram-negative — ensuring HMW DNA suitable for multi-kb read lengths regardless of taxonomic group.
Genomic DNA is extracted using optimized protocols that preserve DNA integrity and minimize shearing. For long-read sequencing, high-molecular-weight (HMW) DNA with fragment lengths exceeding 20 kb is essential for maximizing read length and assembly contiguity. DNA quality is assessed by agarose gel electrophoresis, Qubit fluorometry, and NanoDrop spectrophotometry. Pulsed-field gel electrophoresis (PFGE) or TapeStation analysis is performed for HMW DNA size verification.
Library preparation is platform-specific and optimized for bacterial genome size and complexity. For PacBio HiFi, 15–20 kb SMRTbell libraries are prepared with barcoded overhang adapters for multiplexing (up to 384 samples per SMRT Cell). For Oxford Nanopore, native or ligation-based libraries (SQK-LSK114) are prepared with native barcoding (up to 96 samples per PromethION flow cell). For Illumina, standard 350 bp paired-end libraries are prepared for hybrid polishing. All libraries undergo AMPure bead purification and size selection.
Sequencing is performed on the selected platform(s) to the target coverage depth. Typical coverage targets: 50–100× for PacBio HiFi CCS, 50–100× for ONT PromethION (raw reads), and 50–100× for Illumina polishing reads. For hybrid projects, all sequencing is coordinated to ensure synchronized data delivery.
Figure 1. Complete bacterial whole-genome de novo sequencing workflow — from HMW DNA extraction and multi-platform library preparation through genome assembly, polishing, structural and functional annotation, to GenBank-compatible submission and publication-ready reporting.
Raw reads are processed through a platform-optimized assembly pipeline. For PacBio HiFi, we use hifiasm or Flye (with HiFi mode) for efficient assembly of high-accuracy reads. For ONT, we use Flye or CANU with R10.4.1 Dorado SUP basecalled reads, followed by Medaka polishing. For hybrid assemblies, Unicycler integrates Illumina and long reads for optimal assembly. Assembly quality is evaluated with QUAST, CheckM, and BUSCO, and circularization is verified by inspecting overlapping contig ends and confirming orientation.
Completed assemblies undergo structural annotation (gene prediction via Prokka or Bakta), functional annotation against 12+ databases, and comprehensive report generation. Deliverables include annotated GenBank files, circular genome maps, variant/feature tables, and a detailed project report.
Our bioinformatics pipeline integrates platform-specific assemblers, polishing tools, and a comprehensive suite of annotation databases to deliver publication-ready bacterial genomes. Analysis features are organized across two service tiers.
| Analysis Feature | Basic Package | Advanced Package |
| De novo genome assembly | ✓ Flye / hifiasm / Unicycler | ✓ + CANU, NextDenovo; multi-assembler consensus |
| Assembly polishing & error correction | ✓ Pilon (Illumina), Medaka (ONT) | ✓ + Polypolish, Racon, HERRO; multi-round polishing |
| Assembly QC & completeness assessment | ✓ QUAST, CheckM, coverage statistics | ✓ + BUSCO, Merqury, K-mer completeness analysis |
| Chromosome circularization & gap closure | ✓ Manual circularization verification; Circlator | ✓ + Long-read spanning analysis; T2T gap closure report |
| Structural annotation (gene prediction) | ✓ Prokka (PGAP-compatible) | ✓ Bakta (with Swiss-Prot/UniProt); RNA/tRNA/rRNA prediction |
| Functional annotation databases | ✓ Nr, GO, COG, KEGG | ✓ + CARD, VFDB, CAZy, TCDB, Pfam, antiSMASH, IMG/MER, CRISPRCasFinder |
| Mobile genetic element analysis | — | ✓ IS element identification, prophage prediction (PHASTER), genomic island detection (IslandViewer), integron analysis (IntegronFinder) |
| AMR & virulence gene profiling | ✓ CARD (AMR) + VFDB (virulence) annotation | ✓ + ResFinder, PlasmidFinder, point mutation screening; chromosomal vs plasmid location assignment |
| Comparative genomics | — | ✓ Pan-genome analysis (Roary/PanX), core/accessory genome SNP phylogeny (IQ-TREE), genome alignment (Mauve/Sibelia) |
| Circular genome visualization | ✓ Static circular map (CGView / DNAPlotter) | ✓ Interactive Circos plot with multi-track annotation (GC content, CDS, RNA, AMR, VF, BGCs) |
| Data delivery & reporting | ✓ Annotated GenBank/EMBL/FASTA, assembly report PDF | ✓ + Interactive genome browser (JBrowse2), publication-ready figures, NCBI submission support |
Each sequencing platform offers distinct trade-offs between read accuracy, read length, throughput, and cost. The optimal choice depends on genome size and complexity, the required assembly quality, project scale, and budget. Our platform-agnostic team can recommend the best single-platform or hybrid strategy for your project.
| Feature | PacBio HiFi (Sequel II / Revio) | Oxford Nanopore (PromethION) | Hybrid (HiFi + Illumina) |
| Read accuracy | Q30+ (>99.9% CCS) | Q14–Q20 (Dorado SUP R10.4.1) | Q40+ (long-read assembly + Illumina polishing) |
| Typical read length | 15–25 kb (CCS mode) | 10–100+ kb (native long reads) | Long reads + 150 bp PE (short reads) |
| Assembly result | Single-contig circular chromosome; Q50+ consensus | Single-contig circular chromosome; Q40+ after polishing | Single-contig circular chromosome; Q60+ gold standard |
| Coverage required (typical) | 50–100× CCS (0.5–1 SMRT Cell per 8 bacterial genomes at 384-plex) | 50–100× raw reads (1 PromethION flow cell per 24–48 genomes at 96-plex) | 40× HiFi + 50× Illumina |
| Per-genome cost (multiplexed) | $$ (medium) | $ (lowest) | $$$ (highest accuracy) |
| Resolution of complex repeats (>10 kb) | Good (up to read length) | Excellent (span very long repeats) | Excellent (hybrid + long reads) |
| Best suited for | Small to medium bacterial genomes (<8 Mb) requiring maximum single-base accuracy; publication-grade genomes; AMR/virulence locus characterization | Large bacterial genomes (>8 Mb); multi-plasmid strains; projects maximizing genome count per budget; rapid screening | Demanding genomes (high G+C, multi-replicon, large (>10 Mb)); reference-quality genomes for taxonomic type strains; regulatory submissions |
Our project scientists provide a free platform consultation to determine the optimal sequencing strategy for your bacterial genome project. Contact us to discuss your specific requirements.
| Category | Requirement | Notes |
| Sample type | High-molecular-weight genomic DNA, bacterial pellet, cultured colony, or glycerol stock | DNA extraction service available for colony-to-DNA processing; please inquire for challenging isolates |
| Minimum input (gDNA) | 200 ng (Illumina); 15 µg (PacBio HiFi); 10 µg (Nanopore) | Lower input may be accepted with library amplification; QC-failure risk increases below recommended amounts |
| DNA quality | OD260/280: 1.8–2.0; OD260/230: ≥ 1.8 | HMW DNA with fragment lengths ≥ 20 kb recommended for long-read library preparation |
| Genome size | Any bacterial genome size (0.5–15+ Mb) | Coverage adjusted proportionally; multi-replicon genomes (multiple chromosomes, large plasmids) supported with advanced assembly |
| Sample purity | No visible contaminants; minimal host or carrier DNA for isolate samples | Axenic cultures strongly recommended for de novo assembly; metagenomic samples accepted with modified pipeline |
| Shipping conditions | gDNA: ice pack (4°C) or dry ice; Bacterial pellet/glycerol stock: dry ice; Colonies: room temperature (short-term, ≤ 48 h) | See our Sample Submission Guidelines for detailed instructions and shipping recommendations |
| QC Parameter | Minimum Standard | Reference Quality Target |
| Number of contigs | ≤ 1 per replicon | 1 (circular chromosome) + 1 per plasmid |
| Consensus accuracy (QV) | QV >40 | QV >50 (HiFi) / QV >60 (hybrid) |
| CheckM completeness | >95% | >99% |
| CheckM contamination | <5% | <1% |
| BUSCO completeness (lineage-specific) | >90% | >98% |
| Gene annotation completion rate | >85% of predicted CDS annotated | >95% with functional assignment |
500+ Bacterial Species Successfully Sequenced
We have delivered complete de novo genome assemblies for over 500 bacterial species spanning the full taxonomic and genomic complexity range — from streamlined genomes of obligate intracellular pathogens (~1 Mb, low G+C) to large, multi-replicon genomes of environmental actinobacteria (>10 Mb, >70% G+C). This breadth of experience translates into robust, species-tailored protocols that minimize optimization time and maximize first-pass success.
True Multi-Platform Independence
Unlike service providers restricted to a single sequencing technology, we operate PacBio Sequel II / Revio, Oxford Nanopore PromethION, and Illumina NovaSeq systems in parallel. Our recommendations are driven purely by the scientific requirements of your project — not by platform availability constraints. When a hybrid strategy yields the best result, we coordinate all platforms from a single point of contact.
End-to-End Project Management
Each bacterial genome project is assigned a dedicated project scientist who manages the complete workflow — from DNA extraction and QC through library preparation, sequencing, assembly, annotation, and final data delivery. You receive regular progress updates with intermediate QC metrics at each stage, eliminating the need to coordinate across multiple vendors.
Publication & Database Submission Support
Our standard deliverables include annotated genome files formatted for NCBI GenBank submission, with all required metadata (BioProject, BioSample, assembly accession fields). We also provide a comprehensive "Methods" section draft describing the sequencing and assembly parameters used, suitable for inclusion in your manuscript.
Soto-Serrano A, Li W, Panah FM, Hui Y, Atienza P, Fomenkov A, Roberts RJ, Deptula P, Krych L. Matching excellence: Oxford Nanopore Technologies' rise to parity with Pacific Biosciences in genome reconstruction of non-model bacterium with high G+C content. Microbial Genomics. 2024;10(11):001316. doi:10.1099/mgen.0.001316.
The dairy bacterium Propionibacterium freudenreichii has a genome with high G+C content (~67%), a characteristic that poses well-known challenges for sequencing and assembly. High G+C templates produce PCR bias, reduced sequencing coverage in GC-rich regions, and increased sequencing errors — making this species a rigorous benchmark for comparing long-read sequencing platforms. Previous generations of Oxford Nanopore Technology (ONT) were considered inferior to PacBio for bacterial genome assembly due to higher per-base error rates. However, the introduction of ONT R10.4.1 flow cells with V14 chemistry and Dorado super-accurate (SUP) basecalling promised to narrow or close this gap.
In this study, Soto-Serrano et al. systematically compared PacBio HiFi CCS, ONT R10.4.1 (with and without the custom BARSEQ method), and Illumina short-read platforms for their ability to reconstruct complete, high-quality P. freudenreichii genomes — including the detection of methylation motifs.
Four P. freudenreichii strains (two type strains and two commercial strains) were sequenced on three platforms: PacBio Sequel IIe (HiFi CCS, 15–20 kb SMRTbell libraries), Oxford Nanopore PromethION (R10.4.1 flow cells, V14 chemistry, Dorado SUP basecalling), and Illumina NovaSeq 6000 (pair-end 150 bp). The ONT data were analyzed with and without the inclusion of shorter (~3 kb) PCR reads generated by the custom BARSEQ method. Assemblies were performed using Flye (long-read) and Unicycler (hybrid), polished with Medaka (ONT) and Pilon (Illumina), and evaluated using QUAST, CheckM, and BUSCO. Methylation motif detection was performed using PacBio SMRT Link and ONT Nanomotif tools.
Figure 2. Comparative analysis of PacBio HiFi and Oxford Nanopore R10.4.1 for complete genome reconstruction of the high G+C bacterium Propionibacterium freudenreichii. ONT-only assemblies combining native long reads with BARSEQ short reads achieved near-perfect genome quality comparable to PacBio HiFi. Adapted from Soto-Serrano et al. (2024), Microbial Genomics, CC BY 4.0.
This study provides direct empirical evidence that Oxford Nanopore R10.4.1 sequencing — particularly when combined with the BARSEQ method — has reached parity with PacBio HiFi for complete bacterial genome reconstruction, even for genomically challenging high G+C templates. The findings validate ONT as a cost-effective, stand-alone strategy for bacterial genome assembly, while confirming that PacBio HiFi remains the gold standard when maximum single-base accuracy is required. For researchers, this means that high-quality bacterial genome assembly is now accessible across a wider range of budgets and project scales than ever before.
CD Genomics provides free project consultation to help determine the optimal sequencing strategy for your bacterial genome project. Contact our scientists to discuss your specific requirements.
De novo sequencing assembles a genome from scratch without a reference, producing a complete genome sequence for novel isolates or strains lacking a high-quality reference. Resequencing aligns reads to an existing reference genome and identifies differences (SNPs, InDels, structural variants). De novo assembly is required when characterizing a new species, completing a gap-free genome, or resolving complex regions not present in the reference. For well-characterized strains with a high-quality reference genome available, resequencing is more cost-effective.
A single-contig assembly means the genome is assembled as one contiguous sequence, but it may not be verified as circular. A "closed" or "complete" genome is a single-contig assembly where the 5′ and 3′ ends overlap and are confirmed to form a circular chromosome. Our pipeline includes automatic circularization detection through end-overlap analysis and manual verification of circularization junctions. For genomes with plasmids, each replicon is assembled as a separate closed circular contig. We report assembly status (circular vs. linear contig) for each replicon in the final assembly report.
Recommended coverage depends on the platform and genome complexity. For PacBio HiFi, 50–100× CCS coverage (approximately 0.5–1 SMRT Cell for 8 multiplexed bacterial genomes) is sufficient for single-contig assembly. For Oxford Nanopore, 50–100× raw read coverage yields complete assemblies after polishing. For hybrid strategies, 40–50× long-read coverage combined with 50× Illumina coverage produces the highest quality assemblies (QV60+). For very large genomes (>8 Mb), high G+C (>70%), or multi-replicon genomes, we may recommend higher coverage. Our project scientists calculate the optimal coverage during the free consultation phase.
Yes, but the approach differs from single-isolate de novo assembly. For defined mixed cultures (2–5 known strains), we can perform co-assembly with strain-resolved binning. For complex metagenomic samples, we recommend our dedicated Metagenomics Sequencing service, which uses metagenome-specific assemblers (metaMDBG, hifiasm-meta) and binning tools to recover metagenome-assembled genomes (MAGs) from the community. For the highest quality results, we recommend isolating individual strains before de novo sequencing whenever possible.
The standard annotation package includes structural gene prediction and functional annotation against Nr (NCBI non-redundant protein database), GO (Gene Ontology), COG (Clusters of Orthologous Groups), and KEGG (Kyoto Encyclopedia of Genes and Genomes). The advanced package adds specialized databases: CARD (Comprehensive Antibiotic Resistance Database), VFDB (Virulence Factor Database), CAZy (Carbohydrate-Active enZymes), TCDB (Transporter Classification Database), Pfam (protein families), antiSMASH (biosynthetic gene clusters), IMG/MER (integrated microbial genomes), and CRISPRCasFinder for CRISPR array detection. Custom database annotation is available upon request.
Plasmids are assembled alongside the chromosome from the same long-read data. Long-read assemblers (Flye, hifiasm, CANU) naturally resolve plasmid sequences as separate circular contigs when sufficient coverage is present. Our pipeline includes dedicated plasmid verification: each circular contig smaller than the chromosome is checked for replication origin sequences, plasmid-specific features, and comparison against plasmid databases (PlasmidFinder). For projects requiring comprehensive plasmid characterization, the Advanced package includes complete multi-replicon resolution with copy number estimation, conjugative transfer system identification, and chromosomal vs. plasmid localization of AMR and virulence genes.
Standard turnaround for a bacterial whole-genome de novo sequencing project is approximately 4–6 weeks from sample receipt to final data delivery. This timeline includes: DNA extraction and QC (3–5 days), library preparation (2–3 days), sequencing (1–3 days for ONT PromethION, 3–5 days for PacBio HiFi), genome assembly and polishing (5–7 days), and annotation and report generation (3–5 days). For urgent projects, expedited service is available with sequencing initiated within 48 hours of sample receipt. Multi-strain projects benefit from multiplexing: 24–96 barcoded genomes can be processed simultaneously, with incremental delivery as each genome is completed.
Yes. We have successfully assembled complete genomes for bacterial species with G+C content exceeding 70%, including actinobacteria (Streptomyces, Mycobacteria), and other high-G+C taxa. High G+C genomes present specific challenges including PCR amplification bias, reduced sequencing coverage in GC-rich regions, and increased error rates. We address these through: (1) optimized HMW DNA extraction protocols that minimize shearing of GC-rich regions; (2) PCR-free or low-cycle library preparation to reduce amplification bias; (3) platform-specific coverage adjustments (higher coverage for ONT to compensate for GC-dependent error profiles); and (4) multi-platform hybrid assembly combining the strengths of PacBio HiFi accuracy with ONT long-read contiguity. As demonstrated in our case study (Soto-Serrano et al. 2024), both PacBio HiFi and ONT R10.4.1 now achieve complete genome closure for high G+C templates.
1. Complete, closed circular genome sequence (FASTA format) with clear annotation of chromosome, plasmid(s), and all extrachromosomal elements — ready for NCBI GenBank submission.
2. Circular genome visualization map (CGView / Circos format) showing CDS features (color-coded by COG category), RNA genes, G+C content, GC skew, and annotated AMR/virulence gene locations.
3. Comprehensive genome report summarizing assembly statistics (total length, N50, L50, number of contigs, BUSCO completeness, CheckM completeness/contamination), annotation statistics (CDS count, RNA count, functional assignments), and key findings (AMR gene inventory, BGC content, prophage regions).
4. Full project report in PDF format documenting all methods, QC metrics, assembly parameters, annotation results, and supporting data — designed for manuscript methods sections and grant reports.
Figure 3. Representative deliverable formats for bacterial WGS de novo projects. Left: circular chromosome map with multi-track annotation. Center: assembly QC metrics and completeness assessment. Right: functional annotation summary and comparative genomics overview. AI-generated representative data.
References