Need one comparable callset from raw reads, alignments, gVCFs, or existing VCFs? Standardize cohort evidence before GWAS and population analysis.
Population and cohort studies often accumulate variant data at different processing stages. One project may include raw FASTQ files, another may arrive as aligned BAM or CRAM files, and a third may contain per-sample gVCFs or VCFs generated by different laboratories. These files are not automatically comparable simply because they can all be opened by common bioinformatics tools.
CD Genomics provides a VCF generation service that converts compatible sequencing and variant inputs into a documented, standardized cohort VCF or BCF dataset for population genomics research. The workflow can cover alignment and small-variant calling from FASTQ, review and recall from BAM/CRAM, joint genotyping from compatible gVCFs, or representation-aware harmonization of existing VCFs. The route is chosen from the evidence retained in the input files, not from the file extension alone.
Already have files from several batches or providers? A technical intake review can identify which samples can enter one cohort callset directly, which require reprocessing, and which should remain separate.
Figure 1: Multiple sequencing and variant-file inputs converge only after reference, allele, sample, and quality harmonization into a traceable cohort callset.
A cohort VCF is a multisample variant dataset in which the samples, sites, alleles, genotypes, and quality fields have been processed under a defined shared framework. BCF is the binary counterpart to VCF and may be preferable for large datasets because it supports the same logical records with more efficient computation and storage.
The value of a cohort callset is not that every input record appears in one large file. Its value is that the same genomic position and allele have the same meaning across samples, missing and reference genotypes can be interpreted within the limits of the source data, and filtering decisions are documented. This comparability is what supports defensible allele-frequency estimates, relatedness checks, population structure analysis, linkage disequilibrium analysis, imputation, GWAS, and rare-variant aggregation.
A conventional single-sample VCF normally records observed variant sites and may omit confidently reference positions, uncovered positions, and filtered non-variant evidence. If a site appears in sample A but not sample B, absence from sample B's VCF does not by itself show whether B is homozygous reference, insufficiently covered, filtered, or never evaluated. A direct merge cannot recover evidence that was not retained.
By contrast, a compatible gVCF contains genotype likelihood or reference-confidence information across variant and non-variant regions. Cohort joint genotyping can evaluate those records together and assign genotypes consistently at sites discovered across the cohort. When only single-sample VCFs are available, CD Genomics can normalize and harmonize the records and create a documented multisample dataset, but this output is described as VCF harmonization or cohort assembly rather than as equivalent to joint genotyping from gVCFs or alignments.
| Starting Input | What Can Be Evaluated | Typical Route | Important Limitation |
| FASTQ | Read quality, library and platform metadata, reference suitability, coverage, and sample identity. | QC, alignment, post-alignment processing, per-sample calling, cohort genotyping, filtering, and normalization. | Reprocessing is most complete but requires a confirmed reference and appropriate sequencing design. |
| BAM/CRAM | Header, read groups, reference sequence, alignment quality, duplicates, coverage, and calling provenance. | Compatibility review, optional realignment or preprocessing, per-sample calling, and cohort genotyping. | CRAM requires the exact reference; incompatible or incomplete alignments may need reconstruction from FASTQ. |
| gVCF | Reference-confidence blocks, genotype likelihood fields, caller/version, ploidy, regions, and reference build. | Compatibility checks, cohort import, joint genotyping, filtering, and normalization. | gVCFs from incompatible callers, reference builds, target regions, or ploidy models should not be combined blindly. |
| Single-sample VCF | Coordinates, alleles, samples, filters, field definitions, and provenance. | Validation, normalization, decomposition where appropriate, site/sample harmonization, and cohort assembly. | Missing records cannot reliably distinguish reference calls from no-calls without supporting evidence. |
| Existing cohort VCF/BCF | Header consistency, sample uniqueness, representation, filters, missingness, and downstream compatibility. | Audit, normalization, filtering, subset generation, reheadering, and analysis-ready export. | A polished file cannot repair absent evidence or undocumented upstream bias. |
Input acceptance is project-dependent. The review also considers organism, reference quality, expected ploidy, sex chromosomes where relevant, target regions, sequencing platform, library design, read length, cohort groups, batches, family structure if available, and intended analyses. Structural variants, small variants, organellar variants, and mixed-ploidy datasets require distinct representation and QC choices and are scoped separately when needed.
Figure 2: Input routing determines whether the cohort should be recalled from reads, jointly genotyped from gVCFs, or assembled with explicit limitations from existing VCF records.
All samples must resolve to an agreed reference sequence and coordinate system. Review includes the assembly accession or version, contig names and order, sequence dictionaries, reference checksums where available, alternate or decoy contigs, target intervals, and chromosome conventions. Coordinate liftover alone does not guarantee equivalent alleles; uncertain or failed projections are excluded or reported.
The same biological small variant can be written in more than one valid way, especially around repeats and multiallelic sites. Normalization may include reference-allele validation, left alignment, parsimonious representation, splitting or recombining multiallelic records when appropriate, duplicate detection, and consistent handling of symbolic alleles. INFO and FORMAT fields are remapped or recalculated only when their definitions support the transformation.
Sample identifiers must remain unique and traceable from intake to delivery. The workflow checks manifests, file headers, read groups, sample order, sex or ploidy metadata where applicable, batch labels, population groups, family relationships if supplied, and duplicates or unexpected relatedness. Changes are recorded in an identifier map rather than silently overwriting source names.
VCF headers define the meaning, number, and type of INFO, FORMAT, FILTER, contig, and other fields. Conflicting definitions are resolved before records are combined. Site annotations and genotype annotations are retained, renamed, recalculated, or removed according to documented rules; values from different pipelines are not treated as interchangeable merely because their field IDs match.
1. Project and provenance review
Confirm the research question, organism, assembly, sample set, sequencing design, variant classes, upstream software, file completeness, data-use conditions, and downstream tools.
2. Integrity and identity checks
Validate file readability, indexes, checksums when available, sample names, header consistency, read groups, reference compatibility, and manifest alignment.
3. Route selection by input type
Choose full read-level processing, alignment-level recall, compatible gVCF joint genotyping, or existing-VCF harmonization. Incompatible groups may be reprocessed or kept as separate strata.
4. Per-sample processing
For FASTQ or alignment inputs, perform the agreed read QC, alignment or alignment review, duplicate handling, coverage assessment, and per-sample small-variant calling with reference-confidence output where supported.
5. Cohort-level genotype processing
Import compatible per-sample evidence, discover the union of candidate sites, and assign cohort genotypes using a method suited to sample number, organism, ploidy, coverage, and caller.
6. Normalization and harmonization
Validate reference alleles, standardize record representation, resolve multiallelic sites, unify headers and sample metadata, and construct indexed VCF/BCF outputs.
7. Sample- and variant-level QC
Evaluate missingness, depth and genotype quality distributions, allele counts, heterozygosity, transition/transversion patterns where informative, singleton and multiallelic behavior, batch effects, population-aware statistics, and outliers.
8. Filtering and downstream subsets
Apply transparent filters appropriate to the design, retain a master callset when agreed, and generate derived biallelic or analysis-specific subsets without losing the connection to the source records.
9. Delivery and handoff
Provide callsets, indexes, manifests, QC tables and figures, exclusions, software and parameter records, and notes for the receiving GWAS, LD, structure, imputation, or rare-variant workflow.
Figure 3: Cohort VCF generation is a controlled sequence of provenance review, sample processing, cohort genotyping, representation harmonization, QC, filtering, and documented delivery.
No single metric proves that a cohort VCF is ready for analysis. QC is interpreted across samples, variants, chromosomes or contigs, batches, populations, and allele-frequency ranges. Thresholds are selected from the study design, sequencing depth, expected diversity, organism biology, and intended statistical model rather than copied from an unrelated cohort.
| QC Layer | Examples of Checks | Decision Supported |
| File and reference | Indexes, header definitions, contig dictionary, reference allele match, and duplicate records. | Whether files can be processed together without coordinate or semantic conflicts. |
| Sample | Call rate, depth, genotype quality, heterozygosity, contamination evidence where available, duplicates, relatedness, batch, and population outliers. | Whether each sample should be retained, reviewed, or excluded. |
| Variant | Missingness, allele count, QUAL or model score, depth distribution, strand or mapping annotations when present, and multiallelic status. | Which sites enter the master and downstream subsets. |
| Cohort | Ti/Tv where biologically informative, allele-frequency spectrum, singleton counts, Hardy-Weinberg behavior when appropriate, and batch differences. | Whether technical effects or representation problems remain after processing. |
| Downstream readiness | Biallelic status, phased or unphased state, dosage requirements, marker density, chromosome naming, and sample order. | Whether the delivered subset matches the receiving analysis. |
QC summaries distinguish observed findings from project-specific thresholds. For example, an unusual heterozygosity distribution may reflect contamination, population structure, inbreeding, ploidy, paralogous mapping, or biology. The result should be investigated in context rather than removed automatically. Similarly, Hardy-Weinberg tests are not appropriate for every population design and should not be applied as a universal quality filter.
Figure 4: An illustrative QC dashboard connects file-level compatibility, sample distributions, variant summaries, batch patterns, and downstream-readiness checks.
| Deliverable | Typical Contents | Why It Matters |
| Master cohort VCF/BCF | Normalized multisample records, agreed annotations and filters, BGZF-compressed VCF or BCF, and index files. | Preserves the traceable cohort callset for reuse and controlled derivation. |
| Analysis-specific subsets | Project-dependent biallelic SNP/indel sets, chromosome or interval subsets, and filtered or unfiltered views. | Allows downstream tools to receive the representation they expect. |
| Sample and variant manifests | Final sample order, identifier map, retained and excluded records, and group and batch labels. | Keeps every file linked to the approved cohort design. |
| QC package | Tables and figures for file, sample, variant, cohort, batch, and frequency-stratified checks. | Shows why samples and sites were retained or removed. |
| Processing record | Assembly, reference files, software versions, commands or parameters, filters, and transformations. | Supports reproducibility, audit, and future cohort expansion. |
| Project report | Methods, workflow summary, key findings, limitations, and downstream recommendations. | Connects technical output to the next research decision. |
Exact deliverables are agreed before analysis. Large cohorts may receive chromosome-partitioned VCF/BCF, sparse or database-backed intermediate formats, or cloud-compatible delivery in addition to the requested final exchange files. These implementation choices do not change the requirement that the exported cohort dataset remain standards-compliant and documented.
Best for: cohorts that require consistent upstream processing, contain files from different pipelines, need a new reference assembly, or lack usable genotype-likelihood information. This route provides the strongest control over alignment, callable evidence, per-sample calling, and cohort genotyping.
Not always necessary for: already compatible gVCFs generated with the same reference, regions, caller model, ploidy assumptions, and field definitions. Reprocessing also cannot compensate for unsuitable sequencing coverage or library design.
Best for: cohorts with reference-confidence records that retain the evidence required to evaluate sites discovered in other samples. It supports cohort expansion more defensibly than merging variant-only records.
Not for: gVCFs with unresolved reference mismatches, incompatible field semantics, disjoint target designs, or caller outputs that cannot be jointly interpreted. These may require stratification or recall.
Best for: standardizing coordinates, alleles, headers, identifiers, filters, and record representation when reads or gVCFs are unavailable, or when a downstream tool requires one consistently formatted dataset.
Not equivalent to: cohort joint genotyping. Variant-only VCFs may not preserve homozygous-reference or no-call evidence at sites absent from an individual file. The deliverable and report therefore state what can and cannot be inferred from missing records.
GWAS and quantitative trait analysis: a consistent sample-by-variant matrix reduces avoidable coordinate, allele, and batch discrepancies before Genome-wide Association Analysis. Association models, covariates, and phenotype QC remain downstream tasks.
Population structure and genetic diversity: normalized markers can feed Population Structure Analysis and Genetic Diversity Analysis after relatedness, LD pruning, missingness, and population-specific filters are defined for the study.
Linkage disequilibrium and haplotype analysis: phased or appropriately prepared cohort variants support Linkage Disequilibrium Analysis. The master VCF can be retained while analysis-specific biallelic subsets are generated for the chosen software.
Genotype imputation and phasing: a harmonized cohort callset provides the coordinate and allele consistency needed before genotype imputation and phasing. Imputation-specific strand checks, reference-panel selection, and post-imputation QC are performed in that downstream service.
Rare-variant research: cohort-aware genotypes, allele counts, callable evidence, and transparent filters support rare-variant aggregation. Rare variant association analysis additionally requires phenotype design, annotation, grouping rules, and statistical models.
Cohort expansion and cross-batch studies: documented processing and a stable reference framework make it possible to evaluate new samples against the existing cohort. New batches are not simply appended; compatibility and batch effects are reviewed before release.
Figure 5: One documented master callset can generate purpose-built subsets for GWAS, LD, population structure, imputation, and rare-variant research without conflating their QC requirements.
A Fast, Reproducible, High-throughput Variant Calling Workflow for Population Genomics
Journal: Molecular Biology and Evolution
Published: 2024 (advance publication in 2023)
The following published study illustrates how reference-aware processing, cohort joint genotyping, filtering, and QC can be standardized across diverse population resequencing datasets.
Mirchandani CD, Shultz AJ, Thomas GWC, et al. A Fast, Reproducible, High-throughput Variant Calling Workflow for Population Genomics. Molecular Biology and Evolution. 2024;41(1):msad270.
Public population resequencing datasets are valuable for comparative and conservation genomics, but differences in reference choice, software, filtering, and QC can create technical variation that complicates reuse. The investigators developed snpArcher to make short-read population variant calling more reproducible across nonmodel organisms.
The workflow accepts paired-end short-read FASTQ data and a reference genome, performs read processing and alignment, calls per-sample variants, imports individual records for cohort joint genotyping, filters the resulting multisample callset, and compiles QC summaries. The authors evaluated it on 26 public whole-genome resequencing datasets: 13 bird, 12 fish, and 1 reptile dataset.
The study generated filtered multisample VCFs and standardized QC outputs across datasets that differed in species, genome size, sample number, and sequencing properties. The workflow also recorded the computational steps in a portable configuration, allowing analyses to be rerun and compared under a common framework.
Figure 6: Original case-study summary of 26 public population datasets processed through reference-aware alignment, cohort genotyping, filtering, and QC as reported by Mirchandani et al.
The published work shows why population VCF generation is an end-to-end cohort process rather than a file-concatenation task. Standardized inputs, reference control, cohort genotype processing, QC, and reproducible records make public or newly generated datasets more suitable for comparison and downstream population analysis.
Trust in a cohort callset comes from decisions that remain visible after the pipeline finishes. CD Genomics combines input triage, reference and allele control, cohort-aware processing, multi-level QC, and downstream handoff so that the delivered VCF/BCF can be understood and reused.
Input-route decisions based on retained evidence: FASTQ, BAM/CRAM, gVCF, and VCF inputs are not treated as interchangeable. The feasibility review identifies whether recall, joint genotyping, harmonization, or stratification is scientifically supportable.
Methods, thresholds, deliverable formats, and cohort-expansion strategy are finalized after technical review. No universal accuracy, sample-capacity, variant-yield, or turnaround claim is applied to every project because these outcomes depend on the source evidence, reference, sequencing design, organism, cohort size, and requested analysis.
Potentially, but not by applying one operation to every file. FASTQ and compatible alignments can be recalled under a shared pipeline; compatible gVCFs can enter joint genotyping; existing VCFs can be normalized and harmonized. If source evidence, reference builds, target regions, or caller semantics are incompatible, some inputs may require reprocessing or separate analysis strata.
Variant-only VCFs often omit non-variant evidence. At a site absent from one sample file, a merge cannot determine whether that sample is homozygous reference, uncovered, filtered, or not evaluated. Direct merging can create a formatted multisample file, but it does not reproduce cohort joint genotyping from gVCFs or alignments.
No. Compatibility depends on reference assembly and sequence, caller and version, reference-confidence representation, ploidy, target regions, contig definitions, sample identity, and required fields. The intake review tests these conditions before importing the cohort.
Coordinate liftover may be possible for some records, but it is not a guaranteed one-to-one conversion. Reference alleles, local sequence context, contig structure, and multiallelic representation must be revalidated. Failed or ambiguous variants are reported, and read-level realignment may be preferable when the original data are available.
The master callset is standards-based, but most downstream analyses require additional subsets and QC, such as biallelic SNP selection, missingness filters, MAF thresholds, relatedness review, LD pruning, or chromosome-name conventions. CD Genomics can prepare agreed downstream-ready files while retaining the unambiguous link to the master dataset.
Yes, when the original reference, pipeline, gVCF semantics, sample manifest, and QC rules are retained. New batches are compatibility-checked and reviewed for batch effects. Depending on the method and cohort design, joint genotyping or derived cohort outputs may need to be regenerated rather than appended.
Yes, subject to reference quality, sequencing design, organism biology, caller support, ploidy model, and downstream objective. Diploid human defaults are not automatically applied to plants, animals, microbes, organelles, mixed-ploidy regions, or pooled samples.
For upstream small-variant detection, see Variant Calling or Whole Genome Resequencing. For downstream use of a finalized callset, explore Genome-wide Association Analysis, Linkage Disequilibrium Analysis, Population Structure Analysis, and Genetic Diversity Analysis.
Next step: prepare the organism, reference assembly, sample count, file types, sequencing platforms, batches, target regions, upstream software, ploidy, and intended downstream analyses. These details allow a technical review of the safest route to a comparable cohort callset.
References
For research use only. Not for use in diagnostic procedures.