Turn pathogen genome collections and structured metadata into quality-controlled phylogenomic, lineage, AMR, temporal, and geographic research evidence.
Pathogen genomic surveillance turns collections of bacterial, viral, fungal, or parasitic genomes into evidence about genetic diversity, lineages, relatedness, geographic spread, temporal change, and research-relevant antimicrobial resistance or virulence markers. Its value comes from analyzing genomes together with sampling metadata, not from sequencing isolates one at a time without a population question.
CD Genomics supports projects from purified DNA or RNA inputs and can also begin with FASTQ files, assemblies, consensus genomes, or variant datasets. Cohort-wide QC, reference or assembly strategy, variant analysis, phylogenomics, population structure, lineage or cluster assignment, AMR/virulence profiling, and temporal or geographic comparisons are configured for the pathogen and research design.
Figure 1: Genomic surveillance links pathogen sequence variation with time, place, host, and source metadata.
| Study Type | Typical Question | Potential Outputs |
| Bacterial surveillance | How are isolates related across sites, hosts, or time? | Genome QC, SNP/indel matrix, phylogeny, clusters, lineages, pan-genome context, AMR and virulence markers |
| Viral surveillance | Which variants or lineages are present and how do frequencies change? | Consensus QC, variants, clade/lineage assignment where available, temporal tree, frequency trajectories |
| Fungal or parasitic genomics | What population structure, diversity, or selection patterns occur among isolates? | Variant sets, structure, differentiation, phylogenomics, and candidate adaptive regions |
| One Health research | Are related genomic groups observed across human, animal, food, or environmental sources? | Source-stratified clusters, trees, distance summaries, metadata integration, and comparative figures |
| Anti-infective research | How do resistance markers, lineages, or target variation differ among collections? | Research-grade marker profiles, lineage context, prevalence tables, and candidate comparisons |
Inputs may include genomic DNA, sequencing reads, assemblies, consensus sequences, VCF files, and structured metadata. The service scope is agreed after assessing organism identity, sample type, genome completeness, contamination, host reads, reference availability, cohort size, and the decisions the analysis must support.
| Surveillance Question | Recommended Path | Key Consideration |
| How are isolates related across sites, hosts, or time? | Genome QC, SNP/indel matrix, phylogenomics, and cluster or lineage assignment | Core-region, recombination, and distance rules are defined pathogen-specifically. |
| Which lineages or variants are present and how do their frequencies change? | Consensus or variant QC plus lineage assignment and temporal trajectories | Consistent collection dates and time windows are required before trend analysis. |
| Do resistance or virulence markers differ among collections? | Versioned database comparison with lineage and prevalence context | Genotype does not equal phenotype; phenotypic data and expert interpretation remain important. |
| Are related genomic groups observed across human, animal, food, or environmental sources? | Source-stratified clustering, trees, and distance summaries | Unbalanced geography or source categories can create apparent clusters. |
| Existing reads, assemblies, consensus genomes, or VCF in hand | Analysis-only entry from the earliest available format | Comparable analyses may require reprocessing for consistent QC and filtering. |
| A single isolate without a comparative collection | Outside this surveillance service | Surveillance value comes from analyzing genomes together with sampling metadata. |
Sampling determines what a genomic surveillance study can infer. A dense collection from one site and a sparse collection from many regions answer different questions even if they contain the same number of genomes. Collection date, location resolution, host or source, sampling method, laboratory batch, and inclusion criteria should be structured before trees or clusters are interpreted.
Figure 2: Sampling and metadata structure determine the comparisons that genomic surveillance can support.
| Parameter | Typical Scope | Review Point | Why It Matters |
| Pathogen and material | Bacterial, viral, fungal, or selected parasitic projects | Organism identity, sample type, genome complexity, contamination, and host reads | Genome structure and biology determine the analysis strategy. |
| Input eligibility and documentation | Purified DNA or RNA, as applicable | Input type, purity, integrity, buffer, and transfer conditions | Submission requirements are confirmed during project review. |
| Input format | gDNA, FASTQ, assemblies, consensus genomes, or VCF | Provenance, batch structure, and input completeness | Entry point defines whether consistent reprocessing is needed. |
| QC criteria | Read quality, coverage, contamination, mixed-sample indicators, and duplicates | Pathogen-appropriate thresholds and retained-sample manifest | Comparable genomes are the foundation of surveillance inference. |
| Reference and assembly strategy | Reference-based or assembly-led processing | Core regions, recombination, mobile elements, and missing data | Strategy must match the diversity and structure of the organism. |
| Analysis modules | Variants, phylogenomics, lineages, clusters, pan-genome, AMR/virulence, temporal and geographic patterns | Metadata integration and database/version fields | Genomic evidence stays connected to sampling context. |
| Deliverables | QC tables, alignments, trees, marker tables, figures, and interpretation report | Documented filters and limits on genomic or epidemiological inference | Explicit boundaries prevent overinterpretation of genomic relatedness. |
1. Project and input review
Pathogen, purified nucleic-acid type, cohort, metadata, input format, transfer requirements, and research objectives are reviewed before acceptance.
2. Sequencing or data intake
Projects may include sequencing or begin from existing reads, assemblies, consensus genomes, or variants. Input provenance and batch structure are recorded.
3. Genome-level quality control
Read quality, coverage, contamination, host content, assembly or consensus completeness, mixed-sample indicators, and duplicate records are assessed under pathogen-appropriate criteria.
4. Variant, assembly, or pan-genome processing
A reference-based or assembly-led strategy is selected according to genome diversity and study goals. Core regions, recombination, mobile elements, and missing data are handled explicitly.
5. Population and phylogenomic analysis
Distances, trees, population structure, lineages, clusters, AMR/virulence features, and temporal or geographic patterns are analyzed with relevant metadata.
6. Review and delivery
Outputs include QC status, methods, reference versions, analysis-ready files, figures, and limits on genomic or epidemiological inference.
Figure 3: Genomic surveillance links input review, data intake, genome QC, and phylogenomic analysis so that comparisons remain interpretable.
Comparable genomes are the foundation of surveillance. QC identifies low coverage, contamination, mixed populations, excessive missing sequence, and reference or assembly artifacts. Variant calls are filtered to a consistent callable space before genetic distances or trees are calculated. Existing data may also be processed through Variant Calling.
Maximum-likelihood or other appropriate phylogenies, distance matrices, clustering, and lineage assignments describe genomic relationships at different scales. Recombination and horizontal transfer are considered where relevant. Broader modules include Evolutionary Tree Analysis and Population Structure Analysis.
For sufficiently diverse microbial collections, gene presence/absence and accessory elements may explain relationships not captured by core SNPs. Explore Pan-Genome Sequencing and Accessory Genome Analysis for related workflows.
Detected genes or mutations can be compared with versioned databases and reported with lineage, prevalence, and evidence fields. Genotype does not always equal phenotype, and absence from a database does not prove susceptibility; phenotypic testing and expert interpretation remain important.
Time-resolved trees, lineage-frequency trajectories, source comparisons, and mapped distributions can reveal turnover and expansion patterns. These are genomic research inferences and do not by themselves establish a transmission event or outbreak.
Figure 4: Representative outputs integrate genome QC, population relationships, and sampling metadata.
| Best For | Not For |
| Academic infectious-disease research, microbial evolution, One Health, veterinary, agricultural, foodborne, anti-infective, vaccine, preclinical, multi-site, and longitudinal genome collections | Isolated genomes without a comparative question, collections lacking essential sampling metadata, or studies whose sampling frame cannot support the intended temporal, geographic, or source comparison |
Whole-genome sequencing-based genetic diversity, transmission dynamics, and drug-resistant mutations in Mycobacterium tuberculosis isolated from extrapulmonary tuberculosis patients in western Ethiopia
Journal: Frontiers in Public Health
Published: 2024
Chekesa B, Singh H, Gonzalez-Juarbe N, et al. Whole-genome sequencing-based genetic diversity, transmission dynamics, and drug-resistant mutations in Mycobacterium tuberculosis isolated from extrapulmonary tuberculosis patients in western Ethiopia. Frontiers in Public Health. 2024;12:1399731.
The study addressed limited genomic information about extrapulmonary tuberculosis isolates in western Ethiopia, focusing on lineage diversity, genomic clustering, and resistance-associated mutations.
Whole-genome sequencing was performed for 96 isolates, of which 89 met the study's quality criteria. The authors analyzed lineage and sub-lineage composition, SNP-based genomic relationships, clustering, and resistance-conferring mutations.
Most retained isolates belonged to Lineage 4, with two sub-lineages predominating. The study reported genomic clusters and resistance-associated mutations, illustrating how lineage, distance, and marker evidence can be combined across an isolate collection.
Figure 5: Original case-study summary of the sequencing cohort, lineage analysis, genomic clustering, and resistance-marker findings reported by Chekesa et al. (2024).
The paper demonstrates how WGS, cohort QC, lineage analysis, genomic clustering, and resistance markers can be integrated. Genomic proximity remains evidence to interpret with epidemiological and phenotypic information, not proof of direct transmission or treatment response.
CD Genomics builds surveillance around a defined sampling frame and a pathogen-appropriate analysis plan, preserving the evidence that connects raw sequence quality to population interpretation.
Flexible entry points: projects can combine sequencing with analysis or begin from existing reads, assemblies, consensus genomes, or variants.
Bacterial, viral, fungal, and selected parasitic projects may be considered. Feasibility depends on the organism, material, biosafety requirements, genome characteristics, reference resources, and requested analyses.
Yes. FASTQ files, assemblies, consensus genomes, or variant files can be reviewed. Comparable analysis may require reprocessing from the earliest available format.
No. Genomic similarity can support a hypothesis of relatedness, but direct transmission requires appropriate temporal, geographic, contact, sampling, and epidemiological evidence.
Versioned AMR gene and mutation profiles can be compared across lineages, sources, locations, or study periods to characterize resistance diversity, select representative strains, investigate target variation, and generate hypotheses for anti-infective or preclinical programs. Phenotypic data can be integrated when available.
Useful fields include sample ID, collection date, location at an approved resolution, host or source, specimen type, study site, laboratory batch, and relevant phenotype or exposure variables.
Related workflows include Whole Genome Resequencing, Pan-Genome Sequencing, Variant Calling, Population Structure Analysis, and Evolutionary Tree Analysis.
References