Bacterial Whole-Genome De Novo Sequencing — Complete Genome Assembly by PacBio & Nanopore

Bacterial Whole-Genome De Novo Sequencing — Complete Genome Assembly by PacBio & Nanopore

High-quality complete bacterial genome assembly, circular chromosome closure, plasmid reconstruction, and comprehensive structural and functional annotation by PacBio HiFi CCS, Oxford Nanopore PromethION, and hybrid Illumina+long-read strategies — for pathogenic bacteria, environmental isolates, probiotics, and industrial strains of any G+C content or genomic complexity

Bacterial Whole-Genome De Novo Sequencing — PacBio and Nanopore dual-platform complete bacterial genome assembly service for gap-free circular chromosomes and plasmid reconstruction

CD Genomics provides bacterial whole-genome de novo sequencing on PacBio HiFi (Q30+ CCS), Oxford Nanopore PromethION, and Illumina platforms, delivering complete, closed circular genomes and plasmids for any bacterial species. Our tri-platform pipeline spans HMW DNA extraction through assembly, polishing, and comprehensive functional annotation against 12+ databases — producing publication-ready genomes suitable for GenBank submission without manual gap closure.

Bacterial whole-genome de novo sequencing constructs a complete genome sequence from raw sequencing data without requiring a pre-existing reference genome. Unlike resequencing — which maps reads to an existing reference and identifies differences — de novo assembly builds the genome from scratch, making it indispensable for characterizing novel bacterial isolates, completing gap-free assemblies of known species, and resolving complex genomic regions such as ribosomal operons, mobilome elements, and large structural variants that are invisible to short-read approaches.

At CD Genomics, we provide bacterial whole-genome de novo sequencing on all three major sequencing platforms — PacBio HiFi (Sequel II / Revio), Oxford Nanopore (PromethION R10.4.1), and Illumina NovaSeq — as well as hybrid strategies that combine the strengths of multiple platforms. For bacterial resequencing applications, see our Bacterial Whole-Genome Resequencing service. Our bioinformatics pipelines deliver assembled, annotated, and QC-validated bacterial genomes with circular chromosomes, complete plasmid sequences, and comprehensive functional annotation ready for publication and GenBank submission.

Why Choose Our Bacterial WGS De Novo Service

Long-Read De Novo Assembly Produces Complete Circular Bacterial Genomes That Short Reads Cannot Close

Bacterial whole-genome de novo sequencing is the process of determining the complete DNA sequence of a bacterial genome through computational assembly of raw sequencing reads into contiguous sequences (contigs), and ultimately into complete circular chromosomes and plasmids, without reliance on a reference genome. The quality of the final assembly is fundamentally determined by the read length and accuracy of the sequencing platform used — which is why long-read technologies have transformed bacterial genomics.

Short-read sequencing (Illumina), while cost-effective and accurate per base, produces reads of only 150–300 bp — far shorter than the repetitive elements, insertion sequences, ribosomal RNA operons, and phage integrations that populate even compact bacterial genomes. These features create assembly gaps that fragment the genome into dozens or hundreds of contigs, obscuring the true genomic architecture and preventing complete genome closure. Long-read sequencing platforms — PacBio HiFi (15–25 kb CCS reads at Q30+) and Oxford Nanopore (reads exceeding 100 kb) — span these repetitive regions in single continuous reads, enabling assemblers to resolve complex genomic structures and produce complete, single-contig circular chromosomes.

At CD Genomics, we have established bacterial whole-genome de novo sequencing workflows across all major platforms, with extensive experience spanning over 500 bacterial species — from well-characterized pathogens (Escherichia coli, Staphylococcus aureus, Mycobacterium tuberculosis, Salmonella enterica) to environmental isolates, extremophiles, probiotics, and industrially significant strains across the full taxonomic breadth of the bacterial domain.

Bacterial WGS De Novo Sequencing Delivers Gap-Free Genomes with Accurate Resolution of Repetitive and Mobile Elements

Scientific Advantages

  • Complete, Gap-Free Genomes

Long reads spanning repetitive elements (rRNA operons, IS elements, phage integrations, homopolymer tracts) enable assemblers to produce single-contig circular chromosomes and closed plasmids without manual gap closure — a task that is frequently impossible with short-read data alone.

  • Accurate Resolution of Repetitive & Mobile Elements

Transposable elements, genomic islands, prophage insertions, and tandem gene duplications — all major drivers of bacterial evolution and adaptation — are correctly assembled rather than collapsed or deleted, preserving the true gene content and organization of the genome.

  • Comprehensive Pan-Genome Representation

Complete assemblies from multiple strains of the same species capture the full gene repertoire — including strain-specific accessory genes, antimicrobial resistance determinants, and virulence factors that may be missed in fragmented draft genomes.

Business & Project Advantages

  • Platform-Neutral Optimization

We operate PacBio, ONT, and Illumina platforms in-house and provide unbiased recommendations — from pure long-read assemblies for the highest contiguity to cost-optimized hybrid strategies for large-scale population studies — matched to your specific quality and budget requirements.

  • Scalable From Single Isolates to Large Cohorts

Whether you need a single reference-quality genome for a novel isolate or high-throughput de novo assembly of 500+ strains for a population genomics study, our multiplexed barcoding strategies and multi-platform capacity scale to meet project demands.

  • Regulatory-Ready Documentation

For industrial and pharmaceutical projects, we provide full traceability documentation including sample handling records, sequencing QC metrics, assembly quality statistics, and annotated genome files compatible with regulatory submission requirements.

Bacterial De Novo Sequencing Supports Pathogen Genomics, AMR Profiling, Natural Product Discovery, and Comparative Genomics

Bacterial de novo sequencing is increasingly applied in One Health AMR surveillance programs, where complete genome assemblies from pathogens, livestock, and environmental isolates are compared to track resistance gene transmission across the human-animal-environment interface — a use case that fragmented draft genomes cannot reliably support.

Clinical & Pathogen Genomics

Environmental & Industrial Microbiology

Probiotic & Food Microbiology

Evolutionary & Comparative Genomics

From HMW DNA Extraction to Annotated Genome — The Bacterial De Novo Sequencing Workflow

Our extraction and library protocols are tailored to bacterial cell wall type — enzymatic lysis for Gram-positive and high-G+C actinobacteria, alkaline lysis for Gram-negative — ensuring HMW DNA suitable for multi-kb read lengths regardless of taxonomic group.

1. High-Molecular-Weight DNA Extraction & QC

Genomic DNA is extracted using optimized protocols that preserve DNA integrity and minimize shearing. For long-read sequencing, high-molecular-weight (HMW) DNA with fragment lengths exceeding 20 kb is essential for maximizing read length and assembly contiguity. DNA quality is assessed by agarose gel electrophoresis, Qubit fluorometry, and NanoDrop spectrophotometry. Pulsed-field gel electrophoresis (PFGE) or TapeStation analysis is performed for HMW DNA size verification.

2. Library Construction

Library preparation is platform-specific and optimized for bacterial genome size and complexity. For PacBio HiFi, 15–20 kb SMRTbell libraries are prepared with barcoded overhang adapters for multiplexing (up to 384 samples per SMRT Cell). For Oxford Nanopore, native or ligation-based libraries (SQK-LSK114) are prepared with native barcoding (up to 96 samples per PromethION flow cell). For Illumina, standard 350 bp paired-end libraries are prepared for hybrid polishing. All libraries undergo AMPure bead purification and size selection.

3. Long-Read Sequencing

Sequencing is performed on the selected platform(s) to the target coverage depth. Typical coverage targets: 50–100× for PacBio HiFi CCS, 50–100× for ONT PromethION (raw reads), and 50–100× for Illumina polishing reads. For hybrid projects, all sequencing is coordinated to ensure synchronized data delivery.

End-to-end bacterial whole-genome de novo sequencing workflow from HMW DNA extraction, PacBio/ONT/Illumina library preparation and sequencing, to genome assembly, polishing, annotation, and publication-ready deliverables Figure 1. Complete bacterial whole-genome de novo sequencing workflow — from HMW DNA extraction and multi-platform library preparation through genome assembly, polishing, structural and functional annotation, to GenBank-compatible submission and publication-ready reporting.

4. Genome Assembly

Raw reads are processed through a platform-optimized assembly pipeline. For PacBio HiFi, we use hifiasm or Flye (with HiFi mode) for efficient assembly of high-accuracy reads. For ONT, we use Flye or CANU with R10.4.1 Dorado SUP basecalled reads, followed by Medaka polishing. For hybrid assemblies, Unicycler integrates Illumina and long reads for optimal assembly. Assembly quality is evaluated with QUAST, CheckM, and BUSCO, and circularization is verified by inspecting overlapping contig ends and confirming orientation.

5. Genome Annotation & Report Generation

Completed assemblies undergo structural annotation (gene prediction via Prokka or Bakta), functional annotation against 12+ databases, and comprehensive report generation. Deliverables include annotated GenBank files, circular genome maps, variant/feature tables, and a detailed project report.

Our Bioinformatics Pipeline Delivers Complete Assembly, Multi-Database Annotation, and Comparative Genomics from Long-Read Data

Our bioinformatics pipeline integrates platform-specific assemblers, polishing tools, and a comprehensive suite of annotation databases to deliver publication-ready bacterial genomes. Analysis features are organized across two service tiers.

Analysis Feature Basic Package Advanced Package
De novo genome assembly ✓ Flye / hifiasm / Unicycler ✓ + CANU, NextDenovo; multi-assembler consensus
Assembly polishing & error correction ✓ Pilon (Illumina), Medaka (ONT) ✓ + Polypolish, Racon, HERRO; multi-round polishing
Assembly QC & completeness assessment ✓ QUAST, CheckM, coverage statistics ✓ + BUSCO, Merqury, K-mer completeness analysis
Chromosome circularization & gap closure ✓ Manual circularization verification; Circlator ✓ + Long-read spanning analysis; T2T gap closure report
Structural annotation (gene prediction) ✓ Prokka (PGAP-compatible) ✓ Bakta (with Swiss-Prot/UniProt); RNA/tRNA/rRNA prediction
Functional annotation databases ✓ Nr, GO, COG, KEGG ✓ + CARD, VFDB, CAZy, TCDB, Pfam, antiSMASH, IMG/MER, CRISPRCasFinder
Mobile genetic element analysis ✓ IS element identification, prophage prediction (PHASTER), genomic island detection (IslandViewer), integron analysis (IntegronFinder)
AMR & virulence gene profiling ✓ CARD (AMR) + VFDB (virulence) annotation ✓ + ResFinder, PlasmidFinder, point mutation screening; chromosomal vs plasmid location assignment
Comparative genomics ✓ Pan-genome analysis (Roary/PanX), core/accessory genome SNP phylogeny (IQ-TREE), genome alignment (Mauve/Sibelia)
Circular genome visualization ✓ Static circular map (CGView / DNAPlotter) ✓ Interactive Circos plot with multi-track annotation (GC content, CDS, RNA, AMR, VF, BGCs)
Data delivery & reporting ✓ Annotated GenBank/EMBL/FASTA, assembly report PDF ✓ + Interactive genome browser (JBrowse2), publication-ready figures, NCBI submission support

PacBio HiFi vs. Oxford Nanopore vs. Hybrid — Platform Selection for Your Bacterial Genome Project

Each sequencing platform offers distinct trade-offs between read accuracy, read length, throughput, and cost. The optimal choice depends on genome size and complexity, the required assembly quality, project scale, and budget. Our platform-agnostic team can recommend the best single-platform or hybrid strategy for your project.

Feature PacBio HiFi (Sequel II / Revio) Oxford Nanopore (PromethION) Hybrid (HiFi + Illumina)
Read accuracy Q30+ (>99.9% CCS) Q14–Q20 (Dorado SUP R10.4.1) Q40+ (long-read assembly + Illumina polishing)
Typical read length 15–25 kb (CCS mode) 10–100+ kb (native long reads) Long reads + 150 bp PE (short reads)
Assembly result Single-contig circular chromosome; Q50+ consensus Single-contig circular chromosome; Q40+ after polishing Single-contig circular chromosome; Q60+ gold standard
Coverage required (typical) 50–100× CCS (0.5–1 SMRT Cell per 8 bacterial genomes at 384-plex) 50–100× raw reads (1 PromethION flow cell per 24–48 genomes at 96-plex) 40× HiFi + 50× Illumina
Per-genome cost (multiplexed) $$ (medium) $ (lowest) $$$ (highest accuracy)
Resolution of complex repeats (>10 kb) Good (up to read length) Excellent (span very long repeats) Excellent (hybrid + long reads)
Best suited for Small to medium bacterial genomes (<8 Mb) requiring maximum single-base accuracy; publication-grade genomes; AMR/virulence locus characterization Large bacterial genomes (>8 Mb); multi-plasmid strains; projects maximizing genome count per budget; rapid screening Demanding genomes (high G+C, multi-replicon, large (>10 Mb)); reference-quality genomes for taxonomic type strains; regulatory submissions

Our project scientists provide a free platform consultation to determine the optimal sequencing strategy for your bacterial genome project. Contact us to discuss your specific requirements.

DNA Input and Quality Requirements for Bacterial WGS De Novo Sequencing

Category Requirement Notes
Sample type High-molecular-weight genomic DNA, bacterial pellet, cultured colony, or glycerol stock DNA extraction service available for colony-to-DNA processing; please inquire for challenging isolates
Minimum input (gDNA) 200 ng (Illumina); 15 µg (PacBio HiFi); 10 µg (Nanopore) Lower input may be accepted with library amplification; QC-failure risk increases below recommended amounts
DNA quality OD260/280: 1.8–2.0; OD260/230: ≥ 1.8 HMW DNA with fragment lengths ≥ 20 kb recommended for long-read library preparation
Genome size Any bacterial genome size (0.5–15+ Mb) Coverage adjusted proportionally; multi-replicon genomes (multiple chromosomes, large plasmids) supported with advanced assembly
Sample purity No visible contaminants; minimal host or carrier DNA for isolate samples Axenic cultures strongly recommended for de novo assembly; metagenomic samples accepted with modified pipeline
Shipping conditions gDNA: ice pack (4°C) or dry ice; Bacterial pellet/glycerol stock: dry ice; Colonies: room temperature (short-term, ≤ 48 h) See our Sample Submission Guidelines for detailed instructions and shipping recommendations

QC Standards and Data Interpretation Boundaries for Bacterial Genome Assembly

Assembly Quality Metrics

QC Parameter Minimum Standard Reference Quality Target
Number of contigs ≤ 1 per replicon 1 (circular chromosome) + 1 per plasmid
Consensus accuracy (QV) QV >40 QV >50 (HiFi) / QV >60 (hybrid)
CheckM completeness >95% >99%
CheckM contamination <5% <1%
BUSCO completeness (lineage-specific) >90% >98%
Gene annotation completion rate >85% of predicted CDS annotated >95% with functional assignment

Interpretation Boundaries

  • Bacterial genome assembly data are for research use only. While our assemblies meet reference-quality standards, all genome sequence data should be validated by orthogonal methods before use in clinical diagnostic or regulatory decision-making contexts
  • Assembly quality is fundamentally limited by input DNA quality. Fragmented or contaminated DNA samples will produce fragmented assemblies regardless of the sequencing platform or bioinformatics pipeline used. HMW DNA with fragment lengths ≥ 20 kb is strongly recommended for optimal long-read assembly results
  • Automated annotation is predictive and requires manual curation. Gene predictions and functional annotations generated by automated pipelines (Prokka, Bakta) are computational predictions. For critical genes (AMR determinants, virulence factors, toxin genes), manual curation of the annotation against the primary literature is recommended
  • Plasmid assembly from mixed isolates requires careful interpretation. In samples containing multiple strains or plasmid variants, assemblers may produce hybrid plasmid sequences that do not represent any single biological entity. Axenic cultures are strongly recommended for unambiguous plasmid reconstruction
  • Genome completeness metrics are estimates. CheckM and BUSCO completeness scores are based on conserved single-copy gene content and may underestimate completeness for highly streamlined or atypical genomes. Manual inspection of assembly continuity and gene content is recommended for critical applications

CD Genomics Has Delivered Complete Genomes for 500+ Bacterial Species Across All Major Platforms

500+ Bacterial Species Successfully Sequenced

We have delivered complete de novo genome assemblies for over 500 bacterial species spanning the full taxonomic and genomic complexity range — from streamlined genomes of obligate intracellular pathogens (~1 Mb, low G+C) to large, multi-replicon genomes of environmental actinobacteria (>10 Mb, >70% G+C). This breadth of experience translates into robust, species-tailored protocols that minimize optimization time and maximize first-pass success.

True Multi-Platform Independence

Unlike service providers restricted to a single sequencing technology, we operate PacBio Sequel II / Revio, Oxford Nanopore PromethION, and Illumina NovaSeq systems in parallel. Our recommendations are driven purely by the scientific requirements of your project — not by platform availability constraints. When a hybrid strategy yields the best result, we coordinate all platforms from a single point of contact.

End-to-End Project Management

Each bacterial genome project is assigned a dedicated project scientist who manages the complete workflow — from DNA extraction and QC through library preparation, sequencing, assembly, annotation, and final data delivery. You receive regular progress updates with intermediate QC metrics at each stage, eliminating the need to coordinate across multiple vendors.

Publication & Database Submission Support

Our standard deliverables include annotated genome files formatted for NCBI GenBank submission, with all required metadata (BioProject, BioSample, assembly accession fields). We also provide a comprehensive "Methods" section draft describing the sequencing and assembly parameters used, suitable for inclusion in your manuscript.

Case Study: Nanopore vs. PacBio for Complete Genome Reconstruction of a High G+C Bacterium

Soto-Serrano A, Li W, Panah FM, Hui Y, Atienza P, Fomenkov A, Roberts RJ, Deptula P, Krych L. Matching excellence: Oxford Nanopore Technologies' rise to parity with Pacific Biosciences in genome reconstruction of non-model bacterium with high G+C content. Microbial Genomics. 2024;10(11):001316. doi:10.1099/mgen.0.001316.

1. Background

The dairy bacterium Propionibacterium freudenreichii has a genome with high G+C content (~67%), a characteristic that poses well-known challenges for sequencing and assembly. High G+C templates produce PCR bias, reduced sequencing coverage in GC-rich regions, and increased sequencing errors — making this species a rigorous benchmark for comparing long-read sequencing platforms. Previous generations of Oxford Nanopore Technology (ONT) were considered inferior to PacBio for bacterial genome assembly due to higher per-base error rates. However, the introduction of ONT R10.4.1 flow cells with V14 chemistry and Dorado super-accurate (SUP) basecalling promised to narrow or close this gap.

In this study, Soto-Serrano et al. systematically compared PacBio HiFi CCS, ONT R10.4.1 (with and without the custom BARSEQ method), and Illumina short-read platforms for their ability to reconstruct complete, high-quality P. freudenreichii genomes — including the detection of methylation motifs.

2. Methods

Four P. freudenreichii strains (two type strains and two commercial strains) were sequenced on three platforms: PacBio Sequel IIe (HiFi CCS, 15–20 kb SMRTbell libraries), Oxford Nanopore PromethION (R10.4.1 flow cells, V14 chemistry, Dorado SUP basecalling), and Illumina NovaSeq 6000 (pair-end 150 bp). The ONT data were analyzed with and without the inclusion of shorter (~3 kb) PCR reads generated by the custom BARSEQ method. Assemblies were performed using Flye (long-read) and Unicycler (hybrid), polished with Medaka (ONT) and Pilon (Illumina), and evaluated using QUAST, CheckM, and BUSCO. Methylation motif detection was performed using PacBio SMRT Link and ONT Nanomotif tools.

3. Results

Case study summary — comparison of PacBio HiFi and Oxford Nanopore R10.4.1 for complete genome assembly of Propionibacterium freudenreichii (high G+C), showing assembly statistics, genome completeness, methylation motif detection, and the custom BARSEQ method workflow Figure 2. Comparative analysis of PacBio HiFi and Oxford Nanopore R10.4.1 for complete genome reconstruction of the high G+C bacterium Propionibacterium freudenreichii. ONT-only assemblies combining native long reads with BARSEQ short reads achieved near-perfect genome quality comparable to PacBio HiFi. Adapted from Soto-Serrano et al. (2024), Microbial Genomics, CC BY 4.0.

Key Findings

4. Conclusions

This study provides direct empirical evidence that Oxford Nanopore R10.4.1 sequencing — particularly when combined with the BARSEQ method — has reached parity with PacBio HiFi for complete bacterial genome reconstruction, even for genomically challenging high G+C templates. The findings validate ONT as a cost-effective, stand-alone strategy for bacterial genome assembly, while confirming that PacBio HiFi remains the gold standard when maximum single-base accuracy is required. For researchers, this means that high-quality bacterial genome assembly is now accessible across a wider range of budgets and project scales than ever before.

When to Choose Bacterial De Novo Sequencing — and When Alternative Approaches May Be More Suitable

Choose bacterial whole-genome de novo sequencing when:

Consider alternative approaches when:

CD Genomics provides free project consultation to help determine the optimal sequencing strategy for your bacterial genome project. Contact our scientists to discuss your specific requirements.

Frequently Asked Questions About Bacterial Whole-Genome De Novo Sequencing

Sample Deliverables for Bacterial WGS De Novo Sequencing Projects

1. Complete, closed circular genome sequence (FASTA format) with clear annotation of chromosome, plasmid(s), and all extrachromosomal elements — ready for NCBI GenBank submission.

2. Circular genome visualization map (CGView / Circos format) showing CDS features (color-coded by COG category), RNA genes, G+C content, GC skew, and annotated AMR/virulence gene locations.

3. Comprehensive genome report summarizing assembly statistics (total length, N50, L50, number of contigs, BUSCO completeness, CheckM completeness/contamination), annotation statistics (CDS count, RNA count, functional assignments), and key findings (AMR gene inventory, BGC content, prophage regions).

4. Full project report in PDF format documenting all methods, QC metrics, assembly parameters, annotation results, and supporting data — designed for manuscript methods sections and grant reports.

Representative deliverables from bacterial whole-genome de novo sequencing projects — circular genome map, assembly quality metrics, functional annotation summary, and comparative genomics results Figure 3. Representative deliverable formats for bacterial WGS de novo projects. Left: circular chromosome map with multi-track annotation. Center: assembly QC metrics and completeness assessment. Right: functional annotation summary and comparative genomics overview. AI-generated representative data.

References

  1. Soto-Serrano A, Li W, Panah FM, Hui Y, Atienza P, Fomenkov A, Roberts RJ, Deptula P, Krych L. Matching excellence: Oxford Nanopore Technologies' rise to parity with Pacific Biosciences in genome reconstruction of non-model bacterium with high G+C content. Microbial Genomics. 2024;10(11):001316. doi:10.1099/mgen.0.001316.
  2. Eisenhofer R, Nesme J, Santos-Bay L, Koziol A, Sorensen SJ, Alberdi A, Aizpurua O. A comparison of short-read, HiFi long-read, and hybrid strategies for genome-resolved metagenomics. Microbiology Spectrum. 2024;12(4):e03590-23. doi:10.1128/spectrum.03590-23.
  3. Goussarov G, Mysara M, Cleenwerck I, Claesen J, Leys N, Vandamme P, Van Houdt R. Benchmarking short-, long- and hybrid-read assemblers for metagenome sequencing of complex microbial communities. Microbiology. 2024;170(6):001469. doi:10.1099/mic.0.001469.
Get Your Instant Quote