Probiotic Candidate Strain Genomic Characterization: From Safety Screening to Regulatory Dossier

Inquiry      >

Infographic showing a probiotic candidate strain genome as a circular chromosome surrounded by characterization modules: species identity (ANI/dDDH), mobile elements (plasmids, prophages), antibiotic resistance genes, virulence factors, and functional traits (bacteriocins, adhesion, bile salt hydrolase).

A probiotic candidate strain isolated from a fermented food or a healthy human donor carries a genome that tells the story—its precise taxonomic identity, its cargo of mobile genetic elements, its resistance and virulence gene repertoire, and the functional pathways that underpin its health-promoting potential. Whole genome sequencing has become the definitive tool for probiotic characterization, replacing traditional biochemical tests and single-gene sequencing. This guide covers the genomic workflow that turns a candidate isolate into a well-documented probiotic strain ready for regulatory review.

Key takeaways

  • 16S rRNA alone is insufficient for probiotic strain identification. Average nucleotide identity (ANI) and digital DNA-DNA hybridization (dDDH) against type strains are the EFSA-mandated standards, achieving species- and strain-level resolution that 16S cannot provide.
  • Mobile genetic elements determine safety risk. Plasmid-borne antibiotic resistance genes flanked by insertion sequences pose the greatest concern. Chromosomal, non-mobile resistance genes in the absence of transfer machinery are generally acceptable.
  • Genome mining reveals probiotic mechanisms. Bacteriocin gene clusters, bile salt hydrolase enzymes, mucus-binding proteins, and stress-tolerance systems are all identifiable in silico before a single functional assay is run.
  • EFSA now requires a complete, closed genome. As of 2024, hybrid long-read plus short-read sequencing is the expected standard for bacteria intentionally added to the food chain, replacing the draft-genome era.
  • Phenotype must confirm genotype. Genomic predictions of antibiotic susceptibility or bacteriocin activity must be validated with in vitro assays; genotype-phenotype discrepancies are common and must be resolved before regulatory submission.

Genomics Redefines Probiotic Safety

Probiotic development has undergone a fundamental shift in the past five years. Where strain characterization once relied on biochemical profiling—carbohydrate fermentation panels, growth temperature ranges, Gram staining—and a 16S rRNA gene sequence for taxonomic placement, regulatory agencies now expect a complete, closed genome with comprehensive in silico safety screening. The 2024 EFSA statement on whole genome sequence analysis of microorganisms intentionally used in the food chain made this explicit: hybrid sequencing, complete genome assembly, and multi-database screening for antimicrobial resistance genes are now the baseline.

This shift is driven by hard-won experience. A 2023 metagenomic survey of eight over-the-counter probiotic products by Aziz and colleagues found that some strains carried antibiotic resistance genes and, in one case, virulence genes. A 2024 shotgun metagenomics study of 50 US probiotic supplements by Gundogdu and colleagues identified between 10 and 56 antibiotic resistance genes per product in Bacillus-based formulations alone. A 2025 investigation of infant probiotic products found that 87% contained unlisted bacterial strains. Genomic characterization is not an academic exercise—it is a quality-control necessity.

Species Identity Beyond 16S

The first question any probiotic dossier must answer is: what is this organism, exactly? 16S rRNA gene sequencing, long the workhorse of bacterial identification, fails at the species level for many probiotic genera. The Lactobacillus casei group—L. casei, L. paracasei, and L. rhamnosus—shares near-identical 16S sequences, making single-gene discrimination impossible. The same holds for the Lactobacillus acidophilus group and several Bifidobacterium species complexes commonly found in commercial probiotics.

The EFSA resolution is unambiguous: average nucleotide identity (ANI) or digital DNA-DNA hybridization (dDDH) against the type strain is mandatory. ANI values above 95% confirm species-level identity; dDDH values above 70% serve the same purpose. ANI must also be calculated against type strains of closely related species, particularly when values hover near the 95% threshold. A borderline ANI of 95.2% to L. paracasei and 94.8% to L. casei does not constitute a clean identification, and regulatory reviewers will not accept it as one. For definitive strain-level tracking—essential for intellectual property protection and quality control—SNP profiling against a reference genome provides the resolution that ANI alone cannot. A 2025 single-cell comparative genomics study demonstrated that SNP-based sub-cluster detection can distinguish genomically distinct subpopulations within a single commercial probiotic product, even when ANI values exceed 99.99%. For laboratories developing probiotic dossiers, bacterial whole genome sequencing with hybrid assembly provides the complete, closed genome that regulatory reviewers require.

Mobile Elements and Antibiotic Resistance

Antibiotic resistance genes in a probiotic genome are not an automatic disqualification. The critical distinction—and the one that regulatory assessors scrutinize most closely—is whether those genes are intrinsic and chromosomally encoded, or acquired and carried on mobile genetic elements. An intrinsic tetracycline resistance gene in a Lactobacillus chromosome with no mobile elements in the surrounding genomic neighborhood is low-risk. That same gene on a conjugative plasmid, flanked by insertion sequences, is a red flag.

Prophages complicate the picture further. Most Lactobacillus and Bifidobacterium genomes carry one or more prophage regions, which can range from intact and inducible to degraded remnants of ancient integration events. A 2024 study of Lacticaseibacillus paracasei NY1301 found that its resident 46-kbp prophage is induced at low levels yet the strain passed a high-dose human safety trial with no adverse effects. Prophage presence alone does not equal risk—what matters is whether the prophage encodes toxins, carries antibiotic resistance genes, or can be induced to a lytic cycle.

Plasmids warrant the most scrutiny. Thai Pediococcus acidilactici strains characterized in 2024 by hybrid long-read and short-read sequencing revealed plasmid-borne tet(M) and erm(B) genes, with tet(M) embedded in a composite arrangement flanked by mobile elements. Although incomplete transfer machinery suggested reduced horizontal gene transfer risk, these strains were flagged as requiring additional phenotypic verification—and those from the same study that lacked plasmid-borne resistance genes entirely were advanced as safer candidates—a finding that reinforces the value of full-length plasmid sequencing in candidate characterization. The screening protocol is straightforward: sequence the complete genome including all extrachromosomal elements with a hybrid approach such as Nanopore-based microbial genome sequencing, search against at least two curated antimicrobial resistance databases (CARD, ResFinder, or NDARO) with EFSA-recommended thresholds of ≥80% identity and ≥70% length coverage, and map every hit to its genomic context.

Virulence, Toxins, and Safety Screening

Genes encoding virulence factors and toxins must be screened with the same rigor applied to antibiotic resistance. The same identity and coverage thresholds apply: ≥80% identity and ≥70% query coverage. But interpretation requires nuance. A gene annotated as a "virulence factor" in a pathogenic Escherichia coli genome may encode a housekeeping function—adhesion, stress survival, biofilm formation—in a commensal Lactobacillus. Functional context matters, and the mere presence of a BLAST hit to a virulence database does not make a probiotic strain pathogenic.

The EFSA framework addresses this through the qualified presumption of safety (QPS) concept. Species with an established history of safe use, and for which the absence of acquired antibiotic resistance determinants can be confirmed genomically, receive QPS status. Novel species without a history of safe use face a full safety assessment: absence of virulence determinants must be demonstrated, not assumed. The 2023 ISAPP consensus statement by Merenstein and colleagues reinforced that the presence of any confirmed toxin gene—hemolysins, enterotoxins, certain cytolysins—should disqualify a candidate unless the gene can be shown to be non-functional.

Functional Traits by Genome Mining

Safety screening answers the "do no harm" question. Functional genome mining answers "what does this strain actually do"—and here genomics excels for probiotic development. Sabino and colleagues published a comprehensive probiogenomic workflow in 2025 that includes a curated database of 243 genes associated with probiotic function.

Bacteriocin gene clusters are among the most sought-after genomic features. Class IIa pediocin-like bacteriocins, class I lantibiotics, and circular bacteriocins each have distinctive genetic architectures that are readily identified by tools such as BAGEL4 and antiSMASH. A 2023 metagenomic study found that 92% of bacteriocins identified in probiotic product genomes were novel, underscoring how much functional novelty remains to be discovered. Bile salt hydrolase (BSH) genes, which deconjugate bile acids for gastrointestinal survival, are identified through conserved pfam domains (PF02275, choloylglycine hydrolase). Mucus-binding proteins and sortase-dependent surface proteins that mediate adhesion to the intestinal epithelium are identifiable through LPXTG motif scanning. Stress-tolerance systems—heat shock proteins, cold shock proteins, oxidative stress response regulons—complete the functional picture and help predict manufacturing and gastrointestinal survival.

Comparative genomics adds another dimension: pangenome analyses of multiple strains within a species reveal which functional traits are core to the species and which are strain-specific acquisitions—information that directly informs intellectual property claims and differentiates a commercial product from competitors.

The EFSA WGS Mandate

As of June 2024, the EFSA FEEDAP Panel's updated statement on whole genome sequence analysis establishes an unambiguous set of requirements for microorganisms intentionally used in the food chain. The table below summarizes what a regulatory-grade probiotic genome looks like in 2026.

Requirement EFSA 2024 Specification
Sequencing technology Long-read sequencing or hybrid (short + long) approach required for bacteria
Read depth ≥30× minimum; ≥100× recommended target
Short-read quality Phred score ≥20 per base
Long-read quality Phred score ≥7 average
Genome assembly Complete, closed genome required for bacteria; de novo assembly preferred
Species identification ANI (>95%) or dDDH (>70%) against type strain; compare with closely related species
AMR gene screening At least two curated databases (CARD, ResFinder, NDARO); ≥80% identity, ≥70% length coverage
Virulence/toxin screening Same identity/coverage thresholds; must assess genomic context and transferability
Contamination check <5% reads assigned to unexpected organisms
Data submission Complete genome FASTA, annotation files (.gbk, .fna, .faa), assembly methodology, database versions with accession dates

The EFSA statement explicitly anticipates updates as technologies evolve. Laboratories should plan for periodic re-analysis of genomic data against updated database versions, since new resistance determinants are continuously added to reference databases.

From Sequencing to Regulatory Dossier

A practical characterization workflow converts raw sequence data into the evidence package that a regulatory reviewer expects. The pipeline can be organized into six stages, each producing specific deliverables that feed directly into the dossier.

Stage Action Deliverable
1 Extract high-molecular-weight DNA from a pure culture using optimized sample preparation protocols; verify purity by streaking and Gram stain Master cell bank documentation; DNA integrity report (TapeStation or PFGE)
2 Sequence with a hybrid approach: long reads (Oxford Nanopore or PacBio HiFi) for contiguity plus short reads (Illumina) for base-level accuracy; target ≥100× coverage for both Raw sequencing data (FASTQ); QC report (FastQC, NanoPlot)
3 Assemble with Trycycler (consensus of Flye, Canu, and Raven long-read assemblies, polished with Medaka and Polypolish using short reads); verify circularity with Bandage Complete, closed genome (FASTA); assembly graph visualization; CheckM completeness report
4 Annotate with Bakta or PGAP; calculate ANI/dDDH against type strains with pyANI or fastANI; screen for AMR genes with CARD and ResFinder; screen for virulence factors with VFDB Annotated genome (.gbk, .fna, .faa); ANI/dDDH report; AMR and virulence screening reports with genomic context maps
5 Mine functional traits with BAGEL4 (bacteriocins), antiSMASH (secondary metabolite clusters), and pfam domain searches (BSH, adhesion, stress tolerance); compare with closest reference genomes Functional gene catalog; bacteriocin cluster maps; comparative genomics report
6 Validate genomic predictions with in vitro assays: MIC testing for antibiotics with resistance gene hits; bacteriocin activity assays; bile tolerance and adhesion assays; confirm absence of biogenic amine production Phenotypic validation report; genotype-phenotype concordance table; final safety dossier

Six-stage probiotic characterization workflow diagram from DNA extraction through hybrid sequencing, genome assembly, annotation and safety screening, functional mining, and phenotypic validation.

The genomic data that drives this workflow also serves a commercial purpose beyond regulatory compliance. A fully sequenced and annotated probiotic genome becomes a proprietary asset: it supports patent claims on strain-specific functional traits, enables PCR-based identity testing for quality control across production batches, and provides the reference sequence against which future stability monitoring can detect any drift in the manufacturing process. Choosing the right sequencing platform for your candidate—our guide to choosing microbial sequencing methods compares short-read, long-read, and hybrid approaches. For companies developing multi-strain formulations, microbial diversity analysis by amplicon sequencing can track individual strain abundance across production lots, providing batch-to-batch consistency data that regulatory reviewers increasingly expect to see.

Complete probiotic genome map showing a circular chromosome with annotated features: bacteriocin gene clusters, bile salt hydrolase, mucus-binding proteins, prophage regions, CRISPR arrays, and a plasmid with annotated replication and mobilization genes.

FAQ

Is 16S rRNA sequencing still useful for probiotic characterization?

16S rRNA serves as an initial screening tool—it quickly confirms the genus and provides a rough phylogenetic placement. But it cannot resolve species within the L. casei group, the L. acidophilus group, or several Bifidobacterium complexes. For a regulatory dossier, 16S alone is insufficient. ANI or dDDH against type strains, derived from complete genome data, is the required standard.

What if my probiotic candidate carries a plasmid with an antibiotic resistance gene?

Evaluate three things: the class of antibiotic the gene confers resistance to (WHO medically important antimicrobials are of greatest concern), the mobility context (are insertion sequences, integrases, or conjugative transfer genes in the neighborhood?), and the phenotypic expression (does MIC testing confirm resistance?). A plasmid-borne resistance gene against a clinically critical antibiotic, flanked by transposase genes, is likely disqualifying. An intrinsic chromosomal gene against an antibiotic not used in human medicine, with no mobile elements nearby, is generally acceptable.

How do I know if a virulence gene hit is a real safety concern?

Many genes that appear in virulence factor databases encode functions—adhesion, biofilm formation, stress survival—that are normal components of commensal physiology. Evaluate each hit by asking: does this gene encode a known toxin (hemolysin, enterotoxin, cytolysin)? Is it part of a functional pathogenicity island or type III/IV secretion system? Is it present in the core genome of closely related commensal species? A hemolysin gene in a Lactobacillus genome is a red flag; a fibronectin-binding protein is likely a normal adhesion factor. Consult the EFSA QPS list and published safety literature when uncertain.

Can I skip long-read sequencing and use only Illumina data?

Not for a regulatory submission in 2026. The EFSA 2024 statement requires hybrid or long-read sequencing for bacteria to achieve a complete, closed genome. Draft genomes assembled from short reads alone cannot resolve repeat regions—rRNA operons, insertion sequences, and plasmid-chromosome junctions—and therefore cannot provide the complete mobilome characterization that safety assessment requires. If budget is a constraint, a single multiplexed MinION flow cell combined with existing Illumina data meets the requirement at minimal additional cost.

What functional traits should I prioritize for genome mining?

Start with the traits that define the probiotic mechanism you intend to claim: bacteriocin gene clusters for antimicrobial activity, bile salt hydrolase for gastrointestinal survival, mucus-binding or fibronectin-binding proteins for epithelial adhesion, and stress-tolerance genes for manufacturing robustness. Then expand to secondary metabolite clusters (antiSMASH), exopolysaccharide biosynthesis loci, and vitamin biosynthesis pathways. The Sabino et al. 2025 database of 243 probiotic-associated genes provides a comprehensive reference for systematic mining. Remember that every in silico prediction must be confirmed by in vitro functional assays before it appears in a regulatory dossier or product claim.

References

  1. EFSA (European Food Safety Authority). EFSA statement on the requirements for whole genome sequence analysis of microorganisms intentionally used in the food chain. EFSA Journal. 2024;22(8):e8912. doi:10.2903/j.efsa.2024.8912
  2. Sabino YNV, Paiva AD, Fonseca BR, Medeiros JD, Machado ABF. Deciphering probiotic potential: a comprehensive guide to probiogenomic analyses. Future Microbiology. 2025;20(7-9):611-622. doi:10.1080/17460913.2025.2492472
  3. Merenstein D, Pot B, Leyer G, et al. Emerging issues in probiotic safety: 2023 perspectives. Gut Microbes. 2023;15(1):2185034. doi:10.1080/19490976.2023.2185034
  4. Peng X, Ed-Dra A, Yue M. Whole genome sequencing for the risk assessment of probiotic lactic acid bacteria. Critical Reviews in Food Science and Nutrition. 2023;63(32):11244-11262. doi:10.1080/10408398.2022.2087174
  5. Aziz G, Zaidi A, O'Sullivan DJ. Insights from metagenome-assembled genomes on the genetic stability and safety of over-the-counter probiotic products. Current Genetics. 2023;69(4-6):213-234. doi:10.1007/s00294-023-01271-5
  6. Gundogdu A, Karis G, Killpartrick A, Ulu-Kilic A, Nalbantoglu OU. A shotgun metagenomics investigation into labeling inaccuracies in widely sold probiotic supplements in the USA. Molecular Nutrition & Food Research. 2024;68(12):e2300780. doi:10.1002/mnfr.202300780
  7. Maeno S, Endo A. Inconsistent identification of Apilactobacillus kunkeei-related strains obtained by well-developed overall genome-related indices. Systematic and Applied Microbiology. 2024;47(6):126559. doi:10.1016/j.syapm.2024.126559
  8. Liang H, Zou Y, Wang M, et al. Efficiently constructing complete genomes with CycloneSEQ to fill gaps in bacterial draft assemblies. GigaByte. 2025;2025:gigabyte154. doi:10.46471/gigabyte.154
* For research purposes only, not intended for clinical diagnosis, treatment, or individual health assessments.
Inquiry
Customer Support & Price Inquiry
  • For research purposes only, not intended for clinical diagnosis, treatment, or individual health assessments.