Microbial Whole-Genome De Novo Sequencing — Complete Long-Read Genome Assembly & Annotation

Microbial Whole-Genome De Novo Sequencing — Complete Long-Read Genome Assembly & Annotation

microbial whole-genome de novo sequencing — closed genome assembly from PacBio HiFi and ONT long-read platforms

Short-read assemblies leave gaps, collapse repeats, and routinely miss plasmids — producing fragmented drafts that undermine downstream comparative genomics, AMR surveillance, and publication-quality analysis. CD Genomics' microbial whole-genome de novo sequencing service solves this by combining PacBio HiFi and Oxford Nanopore (ONT) long-read sequencing to deliver closed, chromosome-level assemblies with all plasmids resolved — not draft genomes, but finished references ready for submission and analysis.

Our dual-platform approach serves academic microbiology labs, public health agencies, and biotech teams requiring complete bacterial, fungal, and viral genome assemblies. With over a decade of experience in microbial genomics with long-read sequencing, we provide end-to-end support — from HMW DNA extraction guidance through assembly, polishing, annotation, and comparative bioinformatics.

What you get with our microbial de novo sequencing service

Why Long-Read Sequencing Transforms Microbial De Novo Assembly

Short-read platforms like Illumina have driven microbial genomics for over a decade, but their fundamental limitation — reads of 150–300 bp — creates persistent assembly challenges. Bacterial genomes typically contain ribosomal RNA operons (5–7 kb), insertion sequence (IS) elements (0.7–2 kb), prophage regions (10–100 kb), and near-identical repeat families that exceed single read length. When an assembler encounters a repeat longer than the read, it cannot determine which copy belongs to which genomic location, producing fragmented assemblies with dozens to hundreds of contigs.

The assembly metrics tell the story. A typical Illumina-only assembly of a 5 Mb bacterial genome yields 50–200 contigs, an N50 of 80–150 kb, and routinely misses one or more plasmids. In contrast, PacBio SMRT sequencing with HiFi reads (10–25 kb, >99.9% accuracy) or ONT ultralong reads (up to 100 kb+) can resolve the same genome to 1–3 contigs — often a single circular chromosome plus individual plasmid contigs. This is not incremental improvement; it is a qualitative shift from draft to finished genome status [1].

With a closed genome, you can confidently map structural variants, characterise complete plasmid sequences, resolve IS element positions, identify prophage integration sites, and perform gene neighbourhood analysis without ambiguity about whether adjacent genes are truly adjacent or separated by an assembly break. For labs publishing microbial genomes, submitting complete assemblies to GenBank satisfies increasingly stringent reviewer requirements that prefer closed reference sequences.

Dual-Platform Strategy — PacBio HiFi vs. ONT for Microbial Genomes

No single sequencing platform is optimal for every microbial genome. GC content, genome size, repeat architecture, plasmid complement, and the need for methylation information all influence which technology — or combination — produces the best result. Our dual-platform capability means we do not force your organism into a single workflow; we recommend the platform strategy based on your genome's characteristics and your research questions.

Feature PacBio HiFi (Revio System) ONT (PromethION)
Read Length 10–25 kb (HiFi reads) 10–100+ kb (ultralong possible)
Consensus Accuracy >99.9% (Q30+) >99% (duplex/sup model); ~95% (simplex)
Best Genome Size 1–15 Mb (bacteria, small fungi) 1–100+ Mb (large fungi, complex genomes)
GC Bias Minimal — uniform coverage across GC extremes Minimal — sequence-composition agnostic
Repeat Resolution Excellent for repeats <20 kb Superior for large repeats (rRNA operons, large IS elements)
Plasmid Recovery Excellent — circular consensus resolves small plasmids Excellent — long reads span entire plasmid sequences
Methylation Detection Yes — kinetic-based 5mC/6mA detection (native) Yes — direct signal-based, no chemical conversion needed
Cost per Genome Moderate Moderate–Low (depends on coverage depth)
Best For Most bacterial genomes; publication-quality references; methylation analysis Large fungal genomes; complex repeat regions; ultralong read requirements; real-time analysis

Platform recommendation guide

  • Typical bacterial isolate (2–8 Mb, moderate GC, 0–3 plasmids): PacBio HiFi delivers the optimal balance of accuracy, completeness, and cost. HiFi reads routinely close most bacterial genomes in a single SMRT Cell.
  • GC-rich organisms (Streptomyces, Mycobacterium, ~65–72% GC): Both platforms perform well; PacBio HiFi is preferred for accuracy but ONT may be selected for cost-sensitive multi-isolate projects.
  • Large repeat-rich genomes (fungi >20 Mb): ONT ultralong reads are recommended to span centromeric repeats, telomeric regions, and large transposable element clusters.
  • Plasmid-heavy genomes (multiple large plasmids, >50 kb each): ONT provides the read length advantage for spanning entire plasmid sequences in single reads, simplifying assembly.
  • Methylation-focused studies: PacBio HiFi provides kinetic-based base modification detection without chemical conversion; ONT provides direct signal-based detection — both methods deliver native modification information.
  • Hybrid approach (PacBio + ONT): For the most challenging genomes, combining PacBio HiFi accuracy with ONT read length yields assemblies that neither platform alone can achieve. We design hybrid strategies for genomes with extreme GC content plus large repeats, or when both maximum accuracy and plasmid completeness are required.

For deeper technical background, see our pages on PacBio SMRT sequencing and Oxford Nanopore sequencing technologies, as well as our full-length plasmid sequencing service for focused plasmid analysis.

Complete Service Workflow — From HMW DNA to Closed Genome

Every microbial de novo project follows a validated six-stage pipeline, from sample receipt to delivery of the finished, annotated genome. Each stage includes defined QC checkpoints; no sample proceeds to the next stage without meeting quality thresholds.

microbial whole-genome de novo sequencing workflow — six-stage pipeline from HMW DNA extraction to annotated genome delivery

Stage 1 — HMW DNA Extraction & Quality Control. High-molecular-weight genomic DNA is extracted using organism-optimized protocols (bead-beating for Gram-positive bacteria, enzymatic lysis for fungi and Gram-negatives, CTAB for polysaccharide-rich samples). QC includes Qubit quantification, NanoDrop purity assessment (A260/280 and A260/230), and HMW integrity verification by pulsed-field gel electrophoresis (PFGE) or TapeStation. Samples not meeting minimum thresholds (≥ 2 µg, ≥ 50 ng/µL, A260/280 1.8–2.0) are flagged, and our team contacts you before proceeding.

Stage 2 — Library Preparation. Platform-specific libraries are constructed: SMRTbell libraries with size selection for PacBio HiFi, or native barcoding and adapter ligation kits for ONT. For ultralong ONT reads, library preparation minimizes shearing to preserve fragments >50 kb. Library QC includes fragment size distribution analysis and quantification.

Stage 3 — Sequencing. PacBio libraries are sequenced on the Revio system in HiFi mode, targeting 100× coverage for bacterial genomes and 50–80× for fungal genomes. ONT libraries are sequenced on PromethION flow cells with coverage targets adjusted for genome size and complexity. Real-time basecalling and quality monitoring ensure data quality throughout the run.

Stage 4 — Assembly. Reads undergo quality filtering and adapter trimming. PacBio HiFi data are assembled using hifiasm or Flye with HiFi-specific parameters. ONT data use Flye with --nano-hq or --nano-raw settings, followed by Medaka or Racon polishing [2]. Hybrid strategies combine ONT scaffolding with PacBio HiFi consensus accuracy.

Stage 5 — Polishing & Circularization. Assembled contigs undergo iterative polishing. Circular contigs are identified and terminal overlaps trimmed to produce closed circular sequences. Trycycler-based consensus approaches may be applied for further accuracy improvement [3]. Each closed molecule is verified by read-mapping validation.

Stage 6 — Annotation & Delivery. Standard annotation includes gene prediction (Prodigal for bacteria, Augustus for fungi), functional annotation against COG, GO, KEGG, and CAZy databases, tRNA/rRNA identification, and CRISPR array detection. All deliverables are packaged with assembly statistics, BUSCO completeness scores, and a methods summary suitable for publication.

Assembly & Bioinformatics — Beyond Base-Pair Accuracy

Assembly quality is not only about contig count — it is about whether every biologically meaningful feature of the genome is captured, correctly ordered, and correctly oriented. Our bioinformatics pipeline is designed for both technical excellence (closed genomes, validated circularization) and biological depth (annotation and comparative analysis that supports publication).

Analysis Module Standard Package Advanced Package
Read QC & Filtering ✓ Quality filtering, adapter trimming, length selection ✓ Same as Standard
Assembly ✓ Flye / hifiasm de novo assembly with platform-specific parameters ✓ Multiple assembler comparison; Trycycler consensus; hybrid assembly
Polishing ✓ Medaka / Racon iterative polishing ✓ Deep polishing with multiple tools; manual circularization verification
Assembly QC ✓ QUAST report, BUSCO completeness, contig count & N50 ✓ Same + read depth plots, coverage uniformity analysis, misassembly detection
Gene Prediction ✓ Prodigal (bacteria) / Augustus (fungi) ✓ Multiple gene callers; manual curation of key loci
Functional Annotation ✓ COG, GO, KEGG pathway mapping ✓ CAZy, CARD, VFDB, MEROPS, specialized databases
rRNA / tRNA / ncRNA ✓ Barrnap, tRNAscan-SE ✓ CRISPR array detection, sRNA prediction
Comparative Genomics ✓ Ortholog clustering, pan-genome analysis, synteny mapping
Phylogenetics ✓ Core-genome SNP phylogeny, ANI/AAI calculation, ML tree construction
AMR / Virulence Profiling ✓ CARD/RGI for AMR genes, VFDB for virulence factors; see our antibiotic resistance gene analysis service
Data Visualization ✓ Circular genome map (CGView), basic assembly statistics ✓ Multi-genome comparison plots, phylogenetic trees, heatmaps, synteny diagrams

Every assembly is benchmarked using BUSCO against the appropriate lineage-specific dataset. Assemblies achieving ≥95% BUSCO completeness with ≤3 contigs meet our internal criteria for "closed genome" status. Assemblies that do not meet this standard trigger an investigation — additional sequencing, alternative assembly parameters, or platform supplementation — before delivery.

Microbial Genomes We Support — Bacteria, Fungi, and Viruses

Our microbial de novo sequencing workflows have been applied across the taxonomic spectrum, from small bacterial genomes (~1 Mb) to large fungal genomes (>100 Mb). Each organism type presents distinct assembly challenges; our protocols and platform recommendations are tailored accordingly.

Bacteria. Bacterial genomes (typically 2–10 Mb) are the most common targets for de novo sequencing, and PacBio HiFi routinely closes them in a single run. We have extensive experience with both model organisms (E. coli, Bacillus, Pseudomonas) and non-model environmental isolates, including GC-rich Actinobacteria, AT-rich Mollicutes, and multi-plasmid Enterobacteriaceae. Our bacterial whole-genome de novo sequencing page provides organism-specific details.

Fungi. Fungal genomes (10–100+ Mb) present challenges beyond bacteria: larger size, higher repeat content (transposable elements, telomeric repeats, centromeres), and more complex gene structures with introns. Long-read sequencing is essential for resolving these features. ONT ultralong reads are often recommended for spanning large repetitive regions, while PacBio HiFi provides the accuracy needed for gene annotation. See our fungal whole-genome de novo sequencing service for detailed protocols.

Viruses. Viral genomes (typically 5–200 kb) benefit from long-read sequencing for resolving genomic termini, repeat regions, and mixed-population quasispecies analysis. Direct sequencing of viral DNA or cDNA without amplification eliminates PCR bias and enables detection of minority variants. Our viral genome de novo sequencing service covers both DNA and RNA viruses with appropriate library preparation strategies.

Sample Preparation & HMW DNA Requirements

The quality of your starting DNA is the single most important determinant of assembly success. Degraded, sheared, or contaminated DNA cannot be rescued by downstream processing — regardless of platform or coverage depth.

Sample Type Requirement Notes
DNA Input ≥ 2 µg HMW gDNA Higher input preferred for QC repetition and optional re-sequencing
DNA Concentration ≥ 50 ng/µL (Qubit fluorometer) Lower concentration may require concentration step, which risks shearing
DNA Purity A260/280 = 1.8–2.0; A260/230 ≥ 2.0 Avoid phenol, ethanol, or carbohydrate contamination
HMW Integrity ≥ 30 kb (PacBio); ≥ 10 kb (ONT) Verified by PFGE or TapeStation; sheared DNA <10 kb will produce fragmented assemblies
Sample Volume ≥ 20 µL Ensure adequate volume for QC aliquots plus library preparation
Preservation Buffer TE buffer (pH 8.0) or nuclease-free water Avoid high EDTA concentrations; avoid Tris-only buffers (nuclease risk)
Shipping Dry ice or ice packs in insulated container Use DNA LoBind tubes; minimize freeze-thaw cycles; include clear labeling

What if my DNA doesn't meet these thresholds?

We recognize that field isolates, environmental samples, and archived specimens may not always yield ideal DNA. In such cases, our team can recommend organism-specific extraction protocols, provide guidance on HMW DNA preparation, or — for borderline samples — perform a small-scale QC sequencing run to assess whether usable data can be recovered before committing to a full project.

Why CD Genomics for Microbial De Novo Sequencing

Dual-Platform Without Platform Bias. We operate both PacBio Revio and ONT PromethION systems in-house, which means we recommend the right platform for your genome — not the one we happen to own. For projects where neither platform alone is sufficient, we design hybrid strategies that combine PacBio accuracy with ONT read length.

Closed-Genome Commitment. We do not consider a draft assembly acceptable. Every project is delivered as a closed, chromosome-level assembly — complete chromosomes circularized and verified, all plasmids resolved, zero gaps. If our initial assembly does not meet these standards, we invest additional sequencing and analysis until it does, at no extra cost to you.

Bioinformatics Depth, Not Just Base Calls. Assembly is where most CROs stop. We continue through polishing, circularization, annotation, and — when requested — comparative genomics, phylogenetics, and functional profiling. The deliverable is not a collection of FASTA files; it is a fully characterized genome ready for publication, GenBank submission, or downstream analysis.

QC Rigor at Every Stage. From HMW DNA integrity verification at intake through BUSCO completeness assessment at delivery, every stage has defined QC thresholds. Samples that fail QC are not silently processed — we contact you, explain the issue, and propose solutions before proceeding.

End-to-End Scientific Partnership. Our team includes molecular biologists, bioinformaticians, and microbiologists who understand the biological context of your project. We do not just run pipelines — we discuss experimental design, help interpret results, and support manuscript preparation. Whether you are sequencing a single clinical isolate or a 500-genome population study, we provide the same level of scientific engagement.

FAQs

Demo Results & Data Visualization

Below are representative examples of the visualization outputs included in every microbial de novo sequencing project. These figures are generated from real assembly and annotation data and are designed for direct inclusion in publications, presentations, or grant applications.

Circular Genome Map

circular microbial genome map showing chromosome with GC content, GC skew, CDS tracks, rRNA, tRNA, and plasmidCGView circular genome map of a bacterial isolate assembled by PacBio HiFi. The outermost ring shows CDS on the forward strand; rings 2 and 3 show GC content and GC skew; inner rings display rRNA and tRNA positions. Plasmid contigs are displayed as separate circular maps in the final report.

Assembly Quality Comparison

assembly quality metrics comparison — contig count, N50, genome completeness for short-read vs long-read microbial assembliesAssembly quality comparison between Illumina-only (short-read), PacBio HiFi, ONT, and hybrid approaches for the same ~5 Mb bacterial genome. Bar charts show contig count (lower is better), N50 (higher is better), and BUSCO completeness (higher is better). The hybrid approach achieves the optimal balance of contiguity and accuracy.

Phylogenetic Tree from Comparative Genomics

maximum-likelihood phylogenetic tree based on core genome SNPs — comparative microbial genomics visualizationMaximum-likelihood phylogenetic tree constructed from core-genome SNP alignment of 24 bacterial isolates. The tree is rooted at midpoint and bootstrap support values are shown at nodes. This analysis is included in the Advanced Bioinformatics Package and is suitable for publication as a supplementary figure.

References

  1. Wenger, A.M., Peluso, P., Rowell, W.J. et al. Accurate circular consensus long-read sequencing improves variant detection and assembly of a human genome. Nature Biotechnology 37, 1155–1162 (2019). DOI: 10.1038/s41587-019-0217-9
  2. Kolmogorov, M., Yuan, J., Lin, Y. & Pevzner, P.A. Assembly of long, error-prone reads using repeat graphs. Nature Biotechnology 37, 540–546 (2019). DOI: 10.1038/s41587-019-0072-8
  3. Wick, R.R., Judd, L.M., Cerdeira, L.T. et al. Trycycler: consensus long-read assemblies for bacterial genomes. Genome Biology 22, 266 (2021). DOI: 10.1186/s13059-021-02483-z

For research use only. Not for use in diagnostic procedures.

Get Your Instant Quote