Transform array, reduced-representation, low-pass WGS, VCF, or PLINK data into a harmonized, phased, quality-controlled dataset for downstream population analysis.
Large population cohorts often begin with incomplete genotype information. SNP arrays measure a fixed marker set, reduced-representation methods sample only part of the genome, and low-pass whole-genome sequencing may leave uncertain genotypes at individual sites. Genotype imputation uses observed markers and haplotype patterns from an appropriate reference panel to infer untyped or uncertain variants. Haplotype phasing estimates which alleles occur together on the same parental chromosome copy.
CD Genomics provides a bioinformatics service for cohorts generated by SNP arrays, RRS, GBS, RAD-seq, ddRAD-seq, or low-pass WGS, as well as externally generated VCF and PLINK datasets. The service connects data harmonization, population review, phasing, reference-panel selection, imputation, post-imputation QC, and association-ready delivery. No new sequencing at CD Genomics is required when suitable genotype data already exist.
Already have array, sequencing, VCF, or PLINK data? A technical review can determine whether the marker density, cohort composition, metadata, and available reference resources support reliable phasing and imputation.
Figure 1: Phasing organizes observed alleles into haplotypes, while imputation uses shared haplotype patterns to infer unobserved genotypes and quantify uncertainty.
At a heterozygous site, an unphased genotype identifies two alleles but does not show which parental chromosome carries each one. Phasing reconstructs ordered haplotypes across neighboring markers. Imputation then compares those observed haplotypes with a denser reference panel and estimates genotypes at sites that were not directly measured.
Imputed genotypes are probabilistic estimates, not replacements for direct observation in every context. Dosage values and genotype probabilities preserve uncertainty for downstream statistical analysis, while variant-level quality measures help determine which inferred sites are suitable for a specific study. The useful output is therefore not the largest possible VCF; it is a phased and filtered dataset whose coordinate system, alleles, panel, and QC thresholds are documented.
| Starting Situation | Potential Value | Critical Review Point |
| SNP array cohort | Extends analysis beyond the array's directly typed markers and supports cross-array harmonization. | Array content, strand orientation, genome build, allele frequency differences, and reference-panel overlap must be checked first. |
| RRS, GBS, RAD-seq, or ddRAD-seq | Reduces missingness and can create a denser common marker set for population or association analysis. | Marker distribution, sample depth, species biology, ploidy, and availability of a representative reference panel determine feasibility. |
| Low-pass WGS | Combines genome-wide read evidence with haplotype information to estimate genotypes at reference-panel sites. | Sequencing depth, genotype-likelihood handling, panel composition, and calibration against directly observed genotypes affect reliability. |
| Multiple cohorts or platforms | Creates a shared set of variants for meta-analysis or combined downstream processing. | Cohorts should normally be harmonized and imputed with consistent logic, then compared for batch- and ancestry-specific quality differences. |
Projects may begin with genotype-array exports, VCF/BCF, PLINK BED/BIM/FAM or PGEN/PVAR/PSAM, genotype likelihoods, or project-dependent sequencing-derived variant data. A sample manifest should identify populations, study groups, batches, platforms, sex where relevant, family relationships if available, and intended downstream analyses. Human data also require a clear data-transfer and permitted-use review.
Pre-imputation QC protects the cohort from errors that imputation cannot repair. The review can include sample call rate, duplicate or related samples, sex consistency where applicable, heterozygosity, population outliers, marker missingness, allele frequency, Hardy-Weinberg checks when appropriate, allele coding, strand ambiguity, duplicate positions, reference/alternate allele alignment, and genome-build confirmation.
Figure 2: Input harmonization places heterogeneous genotype sources into a consistent coordinate and allele framework before phasing or imputation.
A file being technically readable does not make it imputation-ready. Sparse exome-only variants, poorly distributed markers, mixed genome builds, unresolved multiallelic sites, severe batch imbalance, or a missing population-relevant reference may limit what can be inferred. The feasibility review records which problems can be corrected, which variants or samples should be excluded, and which limitations remain after processing.
Reference-panel selection is a scientific decision because imputation depends on shared haplotypes between the study cohort and the reference. Panel size and marker density matter, but ancestry representation, genetic similarity, variant frequencies, sequencing quality, genome build, phasing quality, and data-use terms can be equally important. A larger panel may yield more variants without producing the best accuracy for every population or allele-frequency range.
| Panel Factor | Question to Resolve | Decision Impact |
| Ancestry and population coverage | Does the panel contain populations genetically similar to all major cohort groups? | Poor representation can reduce accuracy, particularly for rare or population-specific variants. |
| Genome build and variant representation | Are coordinates, reference alleles, contig names, and multiallelic records compatible? | Unresolved mismatches can remove informative markers or introduce systematic allele errors. |
| Panel size and haplotype diversity | Does the panel provide enough relevant haplotypes for the target population and frequency range? | More haplotypes can increase yield, but irrelevant composition does not guarantee higher accuracy. |
| Target application | Is the priority common-variant GWAS, rare-variant exploration, cohort merging, population analysis, or breeding? | The intended analysis determines which variant classes and quality thresholds deserve emphasis. |
| Access and permitted use | Can the panel legally and operationally be used for the species, cohort, and project type? | Data-use and server terms can constrain panel availability or require an alternative workflow. |
Figure 3: Reference-panel selection balances population fit, genomic compatibility, expected variant spectrum, and permitted use rather than relying on panel size alone.
For human cohorts, public resources such as 1000 Genomes, HRC, TOPMed, or population-focused panels may be considered when access and project scope permit. For plants, animals, or underrepresented human populations, a suitable public panel may not exist. A project-specific panel can be evaluated when phased high-quality genotypes or sequence data are available from representative individuals, but panel construction is an advanced module rather than an automatic component of every project.
When no defensible reference panel is available, the responsible outcome may be to limit imputation, use within-cohort methods appropriate to the species and design, generate additional reference data, or retain the directly observed marker set. CD Genomics does not treat unavailable population representation as a performance problem that software alone can solve.
1. Study, data, and metadata intake
We review the species, genome build, cohort structure, genotyping or sequencing method, file formats, sample manifest, reference resources, and downstream objective. This defines whether the project is a single-cohort imputation, cross-platform harmonization, or an advanced reference-panel task.
2. Sample- and variant-level QC
Samples and markers are evaluated using criteria appropriate to the study design. Exclusions and unresolved issues are recorded so that cohort composition after QC remains traceable.
3. Build, allele, and marker harmonization
Coordinates, chromosome naming, reference alleles, strand orientation, variant identifiers, and multiallelic representation are aligned with the selected reference. Ambiguous or inconsistent sites are filtered or documented rather than silently forced into the analysis.
4. Population and reference-panel assessment
Population structure, ancestry composition, marker overlap, and available reference haplotypes are reviewed. Where useful, candidate panels or stratified workflows can be compared before the final analysis is selected.
5. Haplotype phasing
Observed genotypes are phased with methods suited to the input type, cohort size, pedigree information if available, and reference resources. The exact software and parameters are selected and recorded for the project.
6. Genotype imputation
Phased haplotypes or sequencing likelihoods are imputed against the agreed reference strategy. Processing may be chromosome- or region-based for scalability, followed by consistency checks and file assembly.
7. Post-imputation QC and delivery
Variant-level quality, allele-frequency behavior, sample consistency, marker yield, and project-specific concordance evidence are reviewed. The filtered dataset, dosage or probability information, QC summaries, methods, and downstream-ready formats are then delivered.
Figure 4: The workflow treats harmonization and post-imputation QC as core analytical stages, not optional cleanup around the imputation step.
Quality filtering should follow the intended analysis. INFO or model-based R² values summarize how well an imputed variant is estimated, but thresholds can behave differently across software, reference panels, and allele-frequency bins. MAF-stratified summaries help reveal whether apparently strong overall performance is driven mainly by common variants.
Where independent truth data or masked directly genotyped sites are available, concordance, dosage correlation, non-reference discordance, or error rates can provide additional calibration. Sample-level outliers, chromosomes with unusual retention, allele-frequency shifts, and population-specific differences should be reviewed before association testing. A single universal cutoff is not imposed without considering the cohort and downstream model.
Figure 5: Post-imputation review combines model-based quality, allele-frequency strata, concordance evidence, and sample-level patterns before variants are released for downstream analysis.
| Deliverable | Typical Contents | How It Supports the Next Step |
| Phased genotype dataset | Phased VCF/BCF or another agreed format, subject to input type and project scope. | Supports haplotype-aware processing, LD analysis, and reference-based imputation. |
| Imputed variant dataset | Imputed VCF/BCF with dosages or genotype probabilities and associated quality fields. | Preserves uncertainty for statistical models and enables quality-based filtering. |
| Association-ready files | Filtered PLINK or VCF outputs, sample and variant identifiers, and agreed covariate alignment when in scope. | Reduces reformatting before GWAS or cohort-level analysis. |
| Pre- and post-imputation QC | Sample exclusions, marker harmonization, quality distributions, MAF-stratified retention, and outlier summaries. | Shows which changes occurred and why the final marker set was retained. |
| Panel and harmonization notes | Reference-panel rationale, genome build, allele handling, software versions, parameters, and unresolved limitations. | Makes the analysis reproducible and easier to audit or extend. |
| Project report | Methods, figures, key QC findings, interpretation boundaries, and a delivery manifest. | Connects technical files with the decisions required for downstream use. |
Exact file types depend on the source data, species, software, panel terms, and downstream workflow. Requested formats should be agreed before processing so identifiers, chromosome conventions, and dosage fields remain compatible with the receiving analysis.
GWAS preparation: a harmonized and quality-filtered imputed dataset can increase marker coverage and align cohorts before association testing. The Genome-wide Association Analysis Service is the downstream service that tests genotype-phenotype associations; imputation itself does not perform or validate those associations.
Cohort harmonization: array batches or studies generated on different marker sets can be processed toward a shared variant set, provided platform, population, and reference differences are evaluated rather than hidden.
Population and haplotype studies: phased data can support Linkage Disequilibrium Analysis and Population Structure Analysis. Because phasing and imputation can affect haplotype-based results, these downstream analyses should retain the relevant quality and method information.
Breeding and non-human population research: imputation may help connect sparse genotyping or low-depth data with denser population resources. Species-specific inheritance, ploidy, breeding design, family structure, and reference-panel availability must be evaluated before a workflow is selected.
Best for: cohorts with distributed genome-wide markers, reliable sample metadata, a confirmed reference assembly, and an appropriate public or project-specific reference resource. It is particularly useful before GWAS, meta-analysis, LD analysis, or population comparison when the starting datasets differ in marker density or missingness.
Not the best choice for: projects with very sparse or highly localized markers, unresolved sample identity, mixed genome builds that cannot be reconciled, no representative reference haplotypes, or a requirement to treat every rare inferred genotype as directly observed. When data generation is still needed, consider SNP Genotyping, Reduced-Representation Sequencing, or Genotyping by Sequencing before planning imputation.
A diverse ancestrally-matched reference panel increases genotype imputation accuracy in a underrepresented population
Journal: Scientific Reports
Published: 2023
The following published study illustrates why panel selection and frequency-stratified validation belong in the imputation plan.
Mauleekoonphairoj J, Tongsima S, Khongphatthanayothin A, et al. A diverse ancestrally-matched reference panel increases genotype imputation accuracy in a underrepresented population. Scientific Reports. 2023;13:12360.
The investigators asked how four public reference panels would perform for a Thai cohort that was sparsely represented in major genomic resources. The study compared 1000 Genomes, HRC, GenomeAsia, and TOPMed rather than assuming that the panel with the most samples or variants would be most accurate.
Genome-wide array data from 412 Thai participants were quality controlled, harmonized, phased, and imputed with the four panels. Chromosome 1 imputed genotypes were compared with genotypes called from high-depth WGS in the same samples. The study assessed genotype yield, Minimac-R², genotype concordance, population structure, allele-frequency strata, and selected association results.
TOPMed produced the largest total number of variant sites, while GenomeAsia achieved the highest cohort median genotype concordance rate at 0.974. Accuracy decreased for rare variants across all four panels. The result demonstrates that variant yield and genotype accuracy answer different questions and that ancestry representation can outweigh panel size for a particular cohort.
Figure 6: Original case-study summary of reference-panel choice, variant yield, and genotype concordance reported by Mauleekoonphairoj et al. (2023).
The practical lesson is to evaluate the reference panel against the actual cohort and intended frequency range. A panel that maximizes the number of inferred sites may not maximize concordance, and rare variants require more cautious interpretation than common variants. When direct validation data are available, they can turn panel selection from an assumption into an evidence-based project decision.
Trust in an imputation project comes from traceable choices before, during, and after the algorithm runs. CD Genomics positions imputation as a coordinated cohort analysis, connecting the biological design with file-level harmonization, reference-panel reasoning, quality review, and downstream delivery.
External-data support: suitable VCF, BCF, PLINK, array, reduced-representation, or low-pass sequencing data can enter the workflow without requiring new sequencing at CD Genomics. Input provenance and limitations are documented rather than obscured.
Software, panel, thresholds, and deliverable formats are finalized only after feasibility review. No universal accuracy, marker-yield, sample-capacity, or turnaround claim is applied to every cohort because those outcomes depend on input density, population fit, reference resources, and the analysis objective.
The choice depends on the species, ancestry and population composition, genome build, input markers, target frequency range, downstream analysis, and permitted access. For a diverse cohort, one panel or one threshold may not fit every subgroup. Candidate panels can be compared using marker overlap, ancestry evidence, quality distributions, and direct-genotype concordance when available.
Sometimes, but cohort composition and reference representation must be reviewed first. A cosmopolitan panel may support multiple groups, while stratified phasing, imputation, or QC may be more defensible when population differences are substantial. The final strategy should avoid using ancestry labels as a substitute for genetic evidence.
Assessment can combine INFO or model-based R², allele-frequency bins, retained marker counts, sample and chromosome patterns, allele-frequency consistency, and concordance with masked or independently genotyped sites. The study objective determines which thresholds and variant classes are appropriate; a single headline quality number is not sufficient.
Yes, when the dataset contains sufficient distributed markers and the genome build, allele representation, sample metadata, and permitted use can be confirmed. New sequencing is not required. A feasibility review identifies format conversion, build alignment, variant normalization, or metadata corrections needed before processing.
No. Rare-variant performance is particularly sensitive to input density, allele frequency, haplotype representation, population match, and reference-panel size. Rare inferred variants should be filtered and interpreted using appropriate uncertainty metrics, and important findings may require direct genotyping or sequencing confirmation.
Phasing can be included as part of the imputation workflow or delivered as a separate agreed output. Available family relationships, pedigrees, or high-confidence phase information should be supplied during intake because they may affect method selection and quality assessment. The design depends on species, cohort structure, and data type.
For new marker generation, explore SNP Genotyping, Reduced-Representation Sequencing, or Genotyping by Sequencing. For downstream interpretation of the completed dataset, use Genome-wide Association Analysis, Linkage Disequilibrium Analysis, or Population Structure Analysis.
Next step: prepare the species, genome build, input format, sample count, cohort groups, genotyping or sequencing platform, available reference resources, and intended downstream analysis. These details allow a technical assessment of marker density, panel fit, harmonization requirements, and deliverable format.
References