Bacterial Whole-Genome Resequencing — Strain-Level Variant Discovery & Population Genomics

Bacterial Whole-Genome Resequencing — Strain-Level Variant Discovery & Population Genomics

SNP, InDel, structural variant (SV), and copy number variation (CNV) detection by alignment-based resequencing on PacBio HiFi, Oxford Nanopore PromethION, and Illumina platforms — for antimicrobial resistance profiling, phylogenetic analysis, population genomics, and comparative genomics across bacterial species of any genomic complexity

Bacterial Whole-Genome Resequencing — reference-based variant detection by PacBio, Nanopore, and Illumina platforms for SNP, InDel, SV, and CNV discovery across bacterial strains

Bacterial whole-genome resequencing (WGS resequencing) determines the genetic differences between a newly sequenced bacterial strain and a pre-existing reference genome of the same or closely related species. Unlike de novo sequencing, which builds a genome from scratch, resequencing maps sequencing reads to a reference genome and identifies all classes of genetic variation — from single nucleotide polymorphisms (SNPs) and small insertions/deletions (InDels) to large structural variants (SVs), copy number variations (CNVs), and gene presence-absence polymorphisms. This reference-based approach is the method of choice when a high-quality reference genome is available for the target species, offering the lowest per-genome cost for population-scale studies while delivering the highest sensitivity for variant detection.

At CD Genomics, we provide bacterial WGS resequencing across all major sequencing platforms — PacBio HiFi (Sequel II / Revio), Oxford Nanopore (PromethION R10.4.1), and Illumina NovaSeq — as well as hybrid strategies that combine platforms for optimal variant detection. Our bioinformatics pipelines are optimized for each platform's read characteristics and tailored to the specific variant types most relevant to your study: whether you need ultra-sensitive SNP detection for phylogenetic epidemiology, complete structural variant resolution for AMR mechanism discovery, or population-scale screening of thousands of strains for genome-wide association studies (GWAS).

Why Choose Our Bacterial WGS Resequencing Service

Introduction

Bacterial whole-genome resequencing is the process of sequencing the genome of a bacterial isolate and comparing the resulting reads to an existing reference genome of the same or a closely related species to identify the complete repertoire of genetic differences between the two. Unlike de novo sequencing — which aims to reconstruct the genome sequence itself — resequencing focuses on differences: the specific mutations, insertions, deletions, rearrangements, and gene content changes that distinguish one strain from another. This approach is central to modern bacterial genomics, underpinning applications from outbreak investigation and antimicrobial resistance surveillance to population biology and evolutionary studies.

The choice of sequencing platform fundamentally determines which classes of variation can be reliably detected. Short-read resequencing (Illumina) provides cost-effective, highly accurate detection of SNPs and small InDels, but systematically fails to resolve variation in repetitive regions, mobile genetic elements, and large structural rearrangements — precisely the regions that drive clinically and evolutionarily important bacterial phenotypes. Long-read resequencing on PacBio HiFi or Oxford Nanopore platforms overcomes these limitations: a single contiguous read spanning a complete IS element, prophage integration, or multi-gene resistance cassette allows the bioinformatics pipeline to unambiguously determine the presence, copy number, and genomic context of these elements, rather than collapsing or mis-mapping them.

At CD Genomics, we have delivered bacterial resequencing projects across hundreds of bacterial species and tens of thousands of isolates, supporting clients in public health, clinical microbiology, pharmaceutical development, and academic research. Our service is built around platform-agnostic project design: we select the optimal sequencing and analysis strategy for your specific research question, budget, and sample volume.

Key Advantages of Bacterial WGS Resequencing

Scientific Advantages

  • Complete SV & Mobile Element Resolution

Long reads (10–100+ kb) span entire insertion sequences, transposons, prophage integrations, and tandem gene duplications — resolving structural variants that are invisible or misassembled in short-read data and enabling accurate determination of resistance gene genomic context (chromosomal vs. plasmid-borne).

  • Strain-Level Discrimination

Full recovery of the accessory genome (gene presence-absence across strains) combined with core-genome SNP analysis provides the maximum possible discriminatory power for strain typing, surpassing core-genome MLST (cgMLST) alone.

  • Unbiased AMR Gene & Virulence Factor Profiling

Long-read resequencing captures full-length resistance gene sequences, resolves multi-gene resistance cassettes, and distinguishes between functional and truncated pseudogenes — eliminating the uncertainty of short-read-based AMR predictions.

Business & Project Advantages

  • Cost-Effective at Scale

Multiplexed barcoding (up to 384 samples per PacBio SMRT Cell, up to 96 per ONT flow cell) and optimized coverage targets make long-read resequencing accessible for population-scale projects, with per-strain costs competitive with or below short-read-only approaches when SV resolution is required.

  • Consistent Data Quality Across Batches

Standardized library preparation, sequencing, and bioinformatics protocols ensure batch-to-batch consistency — critical for large population studies and longitudinal surveillance projects where comparability across timepoints is essential.

  • Integrated Multi-Omic Analysis

Long-read resequencing data can simultaneously provide genome sequence, methylation profiles (PacBio HiFi or ONT R10.4.1), and plasmid structures — enabling integrated genomic and epigenomic analysis from a single sequencing run.

Applications of Bacterial Whole-Genome Resequencing

Clinical Microbiology & Public Health

Antimicrobial Resistance Mechanism Discovery

Population & Evolutionary Genomics

Industrial & Food Microbiology

Technology Overview — Bacterial WGS Resequencing Workflow

1. DNA Extraction & QC

Genomic DNA is extracted and purified using standardized protocols appropriate for the bacterial species. For long-read resequencing, high-molecular-weight DNA (fragments ≥ 20 kb) is strongly recommended to maximize read length and mapping accuracy across repetitive regions. DNA quality is verified by agarose gel electrophoresis, Qubit fluorometry, NanoDrop spectrophotometry, and — for HMW DNA — pulsed-field gel electrophoresis (PFGE) or TapeStation fragment analysis.

2. Library Preparation & Barcoding

Platform-specific libraries are prepared with barcoded adapters for multiplexed sequencing. For PacBio HiFi, 15–20 kb SMRTbell libraries with up to 384-sample barcoding. For Oxford Nanopore, native ligation libraries (SQK-LSK114) with native barcoding (up to 96-plex). For Illumina, standard 350 bp paired-end libraries with dual-index barcoding. Libraries are quantified, normalized, and pooled for sequencing.

3. Sequencing

Sequencing is performed to target coverage appropriate for the variant classes of interest. Typical targets: 30–50× for Illumina SNP detection, 30–50× HiFi CCS for PacBio SNP+SV detection, 40–80× raw coverage for ONT SV+SNP detection.

End-to-end bacterial whole-genome resequencing workflow from HMW DNA extraction and multi-platform library preparation through alignment-based variant calling, structural variant detection, and population genomics analysis Figure 1. Complete bacterial whole-genome resequencing workflow — from HMW DNA extraction and multi-platform sequencing through reference alignment, variant detection (SNPs, InDels, SVs, CNVs), functional annotation, and population genomics analysis.

4. Read Alignment & Variant Calling

Quality-filtered reads are aligned to the reference genome using platform-appropriate aligners. For PacBio HiFi: pbmm2 (minimap2-based) or NGMLR. For ONT: minimap2 (with ont preset) or Winnowmap2. For Illumina: BWA-MEM. Variant calling is performed with platform-specific callers: DeepVariant or Clair3 (PacBio HiFi), NanoCaller or Clair3 (ONT), and GATK HaplotypeCaller (Illumina). Structural variants are identified using Sniffles2, DeBreak, cuteSV, or SVDSS for long reads, and Manta or DELLY for Illumina.

5. Annotation & Report Generation

Called variants are annotated for functional impact (SnpEff, ANNOVAR), gene-level consequences, and population genetics metrics. Custom reports summarize variant counts by type, functional category, and genomic region, with visualizations including genome-wide variant density plots, Manhattan plots, and phylogenetic trees.

Bioinformatics Analysis — Variant Detection & Interpretation

Our bioinformatics pipeline supports comprehensive variant detection across all sequencing platforms, with analysis features organized across two service tiers.

Analysis Feature Basic Package Advanced Package
Read QC, alignment & coverage analysis ✓ FastQC, MultiQC, alignment statistics (coverage depth, mapping rate, insert size) ✓ + Platform-specific error profiling, base quality score recalibration (BQSR)
SNP & InDel calling ✓ DeepVariant (HiFi), Clair3 (ONT), GATK (Illumina); VCF output ✓ + Multi-caller consensus; validation against reference calls
Structural variant (SV) detection ✓ Sniffles2 (long reads); Manta (Illumina) ✓ + DeBreak, cuteSV, SVDSS multi-caller integration; SV breakpoint refinement
CNV & gene presence-absence analysis ✓ CNVnator / Smoove (read-depth based); PanX / Panaroo gene content matrix
Variant annotation & functional impact ✓ SnpEff (variant effects: synonymous/non-synonymous, stop-gain, frameshift, etc.) ✓ + ANNOVAR, VEP; custom database annotation; regulatory region annotation
AMR & virulence gene profiling ✓ CARD RGI + VFDB annotation of AMR/VF genes in assembled genome ✓ + ResFinder, PointFinder, PlasmidFinder; chromosomal vs. plasmid location assignment; IS element-mediated resistance mapping
Phylogenetic analysis ✓ Core-genome SNP alignment; IQ-TREE maximum likelihood phylogeny ✓ + Recombination detection (Gubbins, ClonalFrameML); temporal phylogenetics (BEAST2); cgMLST allele-based tree
Population genomics statistics ✓ Nucleotide diversity (π), Tajima's D, FST, PCA, ADMIXTURE ancestry estimation
Pan-genome analysis ✓ Core/accessory genome partitioning (PanX, Panaroo, Roary); gene presence-absence matrix; GWAS association testing (Scoary, pyseer)
Mobile genetic element profiling ✓ IS element mapping, prophage prediction (PHASTER), plasmid typing (PlasmidFinder, MOB-suite), genomic island detection
Custom reporting & visualization ✓ Standard variant report PDF with tables and summary statistics ✓ Interactive genome browser (IGV / JBrowse2), Circos SV plots, publication-ready figures, NCBI BioProject submission files

Choosing the Right Platform for Bacterial Resequencing

The optimal sequencing platform depends on the variant classes you need to detect, genome complexity, sample number, and budget. Our team provides platform-neutral recommendations tailored to each project.

Feature PacBio HiFi Oxford Nanopore Illumina (Short-Read)
SNP detection accuracy ★★★★‍★ (Q30+ HiFi) ★★★‍☆☆ (Q14–Q20, high depth needed) ★★★★‍★ (Q30+ best for SNPs)
InDel detection accuracy ★★★★‍★ (excellent for small InDels) ★★★‍☆☆ (moderate; homopolymer errors) ★★★★‍☆ (good; struggles with >50 bp InDels)
Structural variant detection (>50 bp) ★★★★‍★ (excellent; reads span most SVs) ★★★★‍★ (excellent; ultra-long reads span large SVs) ★★☆☆☆ (poor; most SVs invisible or misassembled)
Mobile element & AMR context resolution ★★★★‍★ (full-length gene cassettes) ★★★★‍★ (complete plasmid sequences) ★★☆☆☆ (fragmented assemblies lose context)
Typical coverage per strain 30–50× CCS 40–80× raw 50–100× PE
Multiplexing (per run) Up to 384 samples Up to 96 samples Up to 384 samples (NovaSeq)
Per-strain cost (multiplexed, 96-plex) $$ (medium) $ (lowest) $$ (medium)
Best suited for SV-focused studies; high-quality variant validation; AMR mechanism discovery; methylome-integrated analysis Large-scale SV screening; plasmid epidemiology; ultra-long repetitive regions; cost-sensitive projects High-throughput SNP typing; core-genome phylogenetics; large GWAS cohorts (>1,000 strains); lowest cost per SNP

Many bacterial resequencing projects benefit from a hybrid strategy: Illumina for high-confidence SNP discovery across the core genome, complemented by long-read sequencing on a subset of representative strains for SV resolution and AMR mechanism characterization. Contact our team for a free project consultation and platform recommendation.

Sample Requirements

Category Requirement Notes
Sample type Genomic DNA, bacterial pellet, cultured colony, or glycerol stock DNA extraction service available; please inquire for challenging or low-biomass samples
Minimum input (gDNA) 500 ng (Illumina); 5 µg (PacBio HiFi); 5 µg (Nanopore) Lower input may be accepted for specific library protocols; QC-failure risk increases below recommended thresholds
DNA quality OD260/280: 1.8–2.0; OD260/230: ≥ 1.8; no visible degradation HMW DNA (≥ 20 kb fragments) recommended for long-read resequencing to maximize SV detection
Reference genome Client-provided or publicly available reference (NCBI RefSeq / GenBank) We can assist with reference selection or recommend the most suitable reference genome for your species
Sample numbers Any scale — from single isolates to large cohorts Batch effects minimized through standardized protocols; large cohort projects receive dedicated project management
Shipping conditions gDNA: ice pack (4°C) or dry ice; Pellet/glycerol stock: dry ice See our Sample Submission Guidelines for detailed instructions

Why Choose CD Genomics for Bacterial Resequencing

Unbiased Platform Expertise

We operate all three major sequencing platforms in-house and base our recommendations purely on the scientific requirements of your project — not on platform availability or commercial agreements. When a novel pathogen requires the highest possible variant resolution, we recommend PacBio HiFi. When cost-efficient population screening is the priority, we guide you toward Illumina or ONT. When a hybrid strategy maximizes value, we coordinate all platforms seamlessly.

Proven Track Record at Scale

Our team has delivered bacterial resequencing projects ranging from single clinical isolates to multi-year population surveillance programs spanning thousands of genomes across dozens of bacterial species. We understand the importance of consistent data quality across batches — essential for longitudinal studies and multi-center collaborations — and maintain strict protocol standardization throughout the project lifecycle.

Specialized AMR & Pathogen Genomics Expertise

Bacterial resequencing for AMR surveillance, outbreak investigation, and pathogen evolution requires domain-specific bioinformatics analysis: resistance gene allele discrimination, plasmid typing, mobile element characterization, and phylogenetic interpretation. Our bioinformatics team has deep expertise in these specialized analyses, with established pipelines for CARD RGI, ResFinder, PlasmidFinder, PointFinder, and custom AMR database annotation.

Dedicated Project Management

Each project receives a dedicated project scientist who manages the complete workflow — sample registration, library preparation, sequencing, bioinformatics analysis, and data delivery — and provides regular progress updates with intermediate QC metrics at each stage.

Case Study: Long-Read Resolves Carbapenemase Alleles Missed by Short-Read Sequencing in Klebsiella pneumoniae

Sierra R, Roch M, Moraz M, Prados J, Vuilleumier N, Emonet S, Andrey DO. Contributions of Long-Read Sequencing for the Detection of Antimicrobial Resistance. Pathogens. 2024;13(9):730. doi:10.3390/pathogens13090730.

1. Background

Antimicrobial resistance (AMR) is one of the most pressing global health threats, and accurate detection of resistance genes is critical for effective patient management, infection control, and epidemiological surveillance. Short-read whole-genome sequencing (WGS) has become the standard tool for AMR gene surveillance, but its limitations in resolving repetitive regions, multi-gene duplications, and complex genomic contexts can lead to incorrect resistance gene allele identification — with potentially serious clinical consequences.

In this study, Sierra et al. performed a head-to-head comparison of Illumina short-read and Oxford Nanopore long-read sequencing for AMR gene detection in a Klebsiella pneumoniae clinical isolate, with a focus on the blaNDM carbapenemase gene family — a group of genes encoding metallo-beta-lactamases that confer resistance to last-resort carbapenem antibiotics.

2. Methods

A K. pneumoniae clinical isolate (VS17) was sequenced on both Illumina (short-read, paired-end) and Oxford Nanopore (ONT R9.4.1, MinION) platforms. Short-read assemblies were generated with SPAdes; long-read assemblies were generated with Flye followed by Medaka polishing and Illumina-based Pilon polishing. AMR gene content was analyzed using ResFinder and CARD RGI. The blaNDM alleles identified by each platform were validated by Sanger sequencing of PCR products spanning the resistance gene locus and by conjugation assays to confirm plasmid transfer and phenotypic resistance profiles.

3. Results

Case study summary — comparison of Illumina short-read vs Oxford Nanopore long-read sequencing for antimicrobial resistance gene detection in Klebsiella pneumoniae, showing NDM carbapenemase allele resolution discrepancy resolved by long reads Figure 2. Comparison of short-read and long-read sequencing for AMR gene detection in Klebsiella pneumoniae. Short-read assembly incorrectly identified a single blaNDM-4 allele, while long-read assembly correctly resolved two distinct alleles (blaNDM-1 and blaNDM-5) located on separate plasmids — confirmed by Sanger sequencing and conjugation assays. Adapted from Sierra et al. (2024), Pathogens, CC BY 4.0.

Key Findings

4. Conclusions

This study provides direct evidence that short-read sequencing alone can produce clinically misleading AMR gene profiles when multiple resistance gene alleles or gene duplications are present. Long-read sequencing resolves these complex loci unambiguously, delivering accurate allele identification and full genomic context. The case underscores the importance of incorporating long-read or hybrid sequencing into clinical AMR surveillance workflows, particularly when accurate resistance gene allele discrimination is required for patient management, outbreak investigation, or molecular epidemiology.

FAQs

Demo

Deliverable Examples for Bacterial WGS Resequencing Projects

1. Variant call file (VCF) with high-confidence SNP, InDel, and SV calls — annotated with functional impact predictions (SnpEff) and platform-specific quality filters (QUAL, GQ, DP).

2. Genome-wide variant landscape visualization — Circos plot showing variant density, SV breakpoints, AMR gene locations, and GC content across the reference genome, with comparative tracks for multiple strains.

3. Phylogenetic tree (Newick format) constructed from core-genome SNP alignments, with bootstrap support values, metadata integration (source, date, phenotype), and publication-ready figure files.

4. Comprehensive project report PDF documenting sample processing, sequencing QC metrics, alignment statistics, variant detection results, annotated gene tables, and detailed methods for manuscript inclusion.

Representative deliverables from bacterial whole-genome resequencing projects — genome-wide variant landscape, phylogenetic tree, AMR gene profiles, and comparative genomics visualizations Figure 3. Representative deliverable formats for bacterial WGS resequencing projects. Left: genome-wide variant density and SV Circos plot. Center: core-genome SNP phylogeny with metadata. Right: AMR gene profile and variant annotation summary. AI-generated representative data.

References

  1. Sierra R, Roch M, Moraz M, Prados J, Vuilleumier N, Emonet S, Andrey DO. Contributions of Long-Read Sequencing for the Detection of Antimicrobial Resistance. Pathogens. 2024;13(9):730. doi:10.3390/pathogens13090730.
  2. O'Donnell S, Yue JX, Abou Saada O, Agier N, Caradec C, Cokelaer T, De Chiara M, Delmas S, Dutreux F, Fournier T, Friedrich A, Kornobis E, Li J, Miao Z, Tattini L, Schacherer J, Liti G, Fischer G. Telomere-to-telomere assemblies of 142 strains characterize the genome structural landscape in Saccharomyces cerevisiae. Nature Genetics. 2023;55:1390–1399. doi:10.1038/s41588-023-01459-y.
  3. Hogle SL, Tamminen M, Hiltunen T. Complete genome sequences of 30 bacterial species from a synthetic community. Microbiology Resource Announcements. 2024;13(6):e00111-24. doi:10.1128/mra.00111-24.
Get Your Instant Quote