Genotype Imputation and Reference Panel Construction for Breeding Populations
Low-density arrays, GBS, LC-WGS, and historical cohorts can leave breeding teams with sparse, missing, or incompatible markers. We design or update a population-matched haplotype reference panel, harmonize and phase target data, validate imputation by masking known genotypes, and deliver versioned VCF and PLINK datasets with clear boundaries for GWAS, genomic selection, and population analysis.
What This Solution Helps You Decide
Genotype Imputation: What It Does
Genotype imputation predicts unobserved variants by matching target samples to haplotypes in a more densely genotyped reference panel. The useful outcome is not simply a larger marker count. It is a harmonized dataset in which every retained dosage is tied to a reference panel version, an imputation-quality rule, and a validation result.
This Solution begins with the downstream decision and works backward. We review the target population, platforms, genome build, marker overlap, and available high-density evidence; then determine whether an existing panel is suitable, a custom panel is needed, or the current panel should be expanded before the next breeding cycle. For the broader breeding context, see Molecular Breeding and Genotyping.
What usually blocks downstream analysis?
What We Can Do with Your Genotype Data
We can assess an existing panel, construct a population-matched reference panel, expand a panel with underrepresented groups, harmonize target cohorts across platforms, and validate where imputed genotypes are reliable enough for downstream analysis.
Reuse is appropriate only when the reference panel contains relevant haplotypes, shares a compatible genome build and marker backbone with the target cohort, and performs acceptably in a masked validation subset. A large panel is not automatically the right panel when the target population is genetically distinct.
Population representation
Do the reference individuals cover the breeds, lines, founders, families, or genetic clusters present in the target cohort? Underrepresented groups may require added high-density samples or a group-specific boundary.
Genome and marker compatibility
Reference and target variants must be reconciled to the same assembly, chromosome naming, coordinates, alleles, and strand convention before phasing or imputation.
Reference genotype quality
High-density or sequence-derived reference genotypes need documented sample QC, variant QC, biallelic representation, missingness handling, and phasing readiness.
Downstream evidence need
The required density and quality threshold depend on whether the dataset will support common-variant GWAS, genomic prediction, population analysis, or evaluation of lower-frequency variants.
A practical panel-fit review produces a decision, not a generic score.
The outcome is to reuse the panel as-is, reuse it with subgroup or variant restrictions, expand it with selected reference individuals, or construct a new version before cohort-wide imputation.
Choose the Reference Panel and Imputation Solution That Fits Your Data
Projects enter at different stages. The starting route determines which data generation, harmonization, phasing, and validation activities are necessary; it does not create artificial service packages.
1. Reuse an existing reference panel
Start here when: a high-density or WGS reference panel already exists for the species and appears to represent the target population.
Work required: audit provenance and version, align the target data, run masking validation, and define subgroup- and MAF-specific release rules.
Decision: whether the existing panel is fit for the planned downstream analysis without additional sequencing.
2. Build a custom breeding-population panel
Start here when: public or legacy panels do not cover the founders, elite parents, regional lines, breeds, or admixed population that must be imputed.
Work required: select representative reference individuals, generate or curate dense genotypes, call and filter variants, phase haplotypes, and validate against held-out truth.
Decision: which reference individuals and panel scope support the current cohort and future cycles.
3. Expand or migrate an operational panel
Start here when: the program has changed platforms, added new germplasm, accumulated new WGS samples, or observed uneven performance across subgroups.
Work required: reconcile panel versions, add qualified haplotypes, rephase where required, repeat masking tests, and document changes from the prior release.
Decision: whether the updated panel improves coverage without breaking cross-cycle comparability.
How Reference Panel Construction and Imputation Work
The workflow builds or qualifies the reference evidence first, aligns the target cohort second, and validates imputation by hiding known genotypes before any dataset is released for GWAS, genomic prediction, or population analysis.
Step 1: Build or Qualify the Reference Panel
Reference-panel construction is a population-design task before it is a software task. The panel must capture relevant haplotypes, use a consistent reference assembly, and preserve enough provenance to be updated and revalidated later.
1. Select representative individuals
Use pedigree, founder status, breed or line labels, population structure, geographic origin, and existing genotype evidence to identify core reference samples and coverage gaps.
2. Generate or curate dense genotypes
Use qualified WGS or high-density genotypes, with route-specific sample and sequencing requirements confirmed before data generation.
3. Apply sample and variant QC
Resolve identity conflicts, duplicates, missingness, build inconsistencies, multiallelic representation, and poorly supported variants before panel assembly.
4. Phase and assemble haplotypes
Select a phasing and imputation strategy that reflects population size, relatedness, marker density, pedigree availability, and computational constraints.
5. Freeze and document the panel version
Record included samples, genome build, variant set, software and parameters, exclusions, and update rules so every imputed cohort remains traceable.
Panel construction connects population coverage to a controlled, versioned haplotype resource; it is not simply a pooled VCF.
Step 2: Align Arrays, GBS, LC-WGS, and WGS Data
Target and reference genotypes must describe the same biological alleles on the same coordinate system. Marker-ID matching alone is not enough when data come from different platforms, genome builds, batches, or variant-calling pipelines.
| Data source | Role in the project | Key alignment need | When it fits | Main boundary |
|---|---|---|---|---|
| SNP arrays | Stable observed backbone for target cohorts | Panel version, marker ID, coordinate, strand, allele, and no-call handling | Repeated breeding cohorts with established species panels | Fixed content may provide weak overlap for distinct populations or new variants |
| GBS | Sequencing-based observed markers for large populations | Reference alignment, locus consistency, missingness pattern, and batch-aware variant representation | Species or populations without a suitable fixed array | Marker presence and depth can vary across samples and batches |
| LC-WGS | Genome-wide genotype likelihoods or sparse calls for target cohorts | Reference build, read and likelihood QC, shared sites, depth distribution, and panel match | Large cohorts requiring genome-wide coverage with reference-supported inference | Performance depends strongly on reference quality and population representation |
| WGS / high-density genotypes | Reference-panel evidence and validation truth | Joint calling or harmonized calls, variant QC, phasing, and consistent build | Core parents, founders, representative lines, breeds, or validation samples | More samples or variants do not compensate for poor representation or unresolved QC |
Confirmed data-generation routes include Whole Genome Sequencing, Low-Coverage WGS, Genotyping by Sequencing, Crop Genotyping Array Services, and Livestock Genotyping Array Services. Route selection is confirmed against the species, population, existing marker backbone, and update plan.
Step 3: Validate Imputation Before Downstream Use
An imputed call should enter GWAS or genomic prediction only after validation shows where the panel performs reliably. We separate model-reported quality scores from empirical masking tests and report performance across the population features that matter to the project.
The analysis sequence: audit target and reference data, harmonize alleles and coordinates, phase compatible genotypes, impute unobserved variants, apply post-imputation QC, and compare imputed calls with genotypes deliberately hidden from qualified validation samples.
What happens when validation is uneven?
The response depends on the failure pattern. We may restrict the released variant set, apply subgroup-specific thresholds, exclude incompatible samples, increase observed marker density, add representative reference individuals, rebuild the panel, or state that the current data should not be used for the planned analysis. Low-confidence values are not promoted as equivalent to observed genotypes.
Results and Deliverables for GWAS and Breeding Analysis
The deliverable package preserves both the inferred genotypes and the evidence needed to decide how they can be used. Exact scope is confirmed during project review.
Reference panel release
A phased, versioned reference panel with sample inclusion, genome build, variant scope, phasing method, and update provenance documented.
Phased and imputed VCF
Target-cohort genotypes or dosages aligned to the agreed assembly, with imputation-quality annotations and retained-variant rules.
PLINK-ready dataset
A downstream analysis set with stable IDs and documented conversion and filtering logic for supported GWAS, population, or breeding workflows.
Masking validation evidence
Concordance, dosage correlation, and quality summaries organized by MAF, subgroup, platform, batch, chromosome, or other project-relevant strata.
QC and exclusion record
Traceable sample and variant decisions, unresolved conflicts, excluded groups, and any route-specific limitations.
Downstream readiness recommendation
A documented decision on which data can proceed, which need filtering or restricted use, and what evidence should be added before the next panel version.
How to Use the Result
A passed dataset can move into the agreed analysis within its validated population and variant boundary. A conditional result may proceed after subgroup or quality filtering. A failed result triggers a defined remediation plan rather than a larger but unreliable genotype matrix. Downstream support is available through GWAS Services and Agricultural Genomic Data Analysis.
Published Research Case: What Controls Imputation Accuracy in Farm Animals
A multi-species study demonstrates why software choice alone cannot define a reliable imputation strategy. Reference-panel size, target-reference relationship, marker density, allele frequency, and parameter settings all changed the evidence available for release.
Published Study
Jiang, Y., Song, H., Gao, H., Zhang, Q., & Ding, X. (2022). Exploring the optimal strategy of imputation from SNP array to whole-genome sequencing data in farm animals. Frontiers in Genetics, 13, 963654. DOI: 10.3389/fgene.2022.963654.
Research question
How do imputation software versions, parameter settings, chip density, reference-panel size, and genetic relationship affect SNP-array-to-WGS imputation in cattle, pigs, and chickens?
Study design
The study analyzed WGS data from 1,682 cattle, 409 pigs, and 335 chickens. Cattle scenarios compared low-, medium-, and high-density marker masks, breed-specific and combined reference groups, multiple Beagle versions, and known genotypes retained as validation truth.
Key findings
Default effective-population-size settings reduced accuracy in small-reference or low-density scenarios. Larger combined reference panels and denser observed marker sets generally improved performance, especially for lower-frequency variants, while genetically distinct target animals showed more variable results.
Why it matters for this solution
The study supports a validation-led strategy: test panel composition, target representation, observed marker density, parameters, and allele-frequency behavior together before deciding that imputed variants are suitable for downstream use.
What this study does not prove
The reported behavior reflects the studied species, populations, panels, software versions, parameters, and validation scenarios. It does not set a universal accuracy threshold or guarantee equivalent performance in another crop, livestock population, platform, or genome build.
Illustration: original conceptual summary based on the cited study. Not a reproduction of the published figure.
Data and Sample Requirements
You can begin with existing genotype data, an existing reference panel, or samples that require new dense genotyping. We can review partial materials first; route-specific data fields and physical sample specifications are confirmed before transfer or shipment.
Target-cohort data
Reference-panel evidence
| Project entry point | What to provide first | What we confirm before execution |
|---|---|---|
| Existing target and reference data | Representative files, sample/variant counts, assembly, platform, population labels, and intended downstream use | Compatibility, reusable evidence, validation subset, compute scope, and release criteria |
| Target data but no suitable panel | Target-platform summary, population structure or pedigree, available founders/core parents, and any public or legacy panel | Reference-sample selection, dense genotyping route, panel coverage, and pilot validation design |
| New genotyping required | Species, population, sample count estimate, sample type, collection status, and project schedule | WGS, LC-WGS, GBS, or array route; extraction needs; route-specific DNA quantity, quality, preservation, and shipping requirements |
When the project should pause
Imputation should not proceed when sample identities are unresolved, the reference and target cannot be aligned to a common assembly, marker overlap is insufficient for phasing, reference genotypes fail QC, or the target population is not represented well enough to support the intended decision. The feasibility review identifies which of these gaps can be repaired.
Why CD Genomics
FAQ
Discuss Your Reference Panel and Imputation Project
Start with a representative file and a short description of the decision the completed dataset must support. We will map the current evidence to a reuse, construction, or update route.
For practical dataset handoff guidance, read Building a GS-Ready Dataset from Array Outputs.
References
Browning, B. L., & Browning, S. R. (2016). Genotype Imputation with Millions of Reference Samples. The American Journal of Human Genetics, 98(1), 116–126. DOI: 10.1016/j.ajhg.2015.11.020.
Das, S., Forer, L., Schönherr, S., Sidore, C., Locke, A. E., Kwong, A., et al. (2016). Next-generation genotype imputation service and methods. Nature Genetics, 48, 1284–1287. DOI: 10.1038/ng.3656.
Yang, W., Yang, Y., Zhao, C., Yang, K., Wang, D., Yang, J., Niu, X., & Gong, J. (2020). Animal-ImputeDB: a comprehensive database with multiple animal reference panels for genotype imputation. Nucleic Acids Research, 48(D1), D659–D667. DOI: 10.1093/nar/gkz854.
Gao, Y., Yang, Z., Yang, W., Yang, Y., Gong, J., Yang, Q.-Y., & Niu, X. (2021). Plant-ImputeDB: an integrated multiple plant reference panel database for genotype imputation. Nucleic Acids Research, 49(D1), D1480–D1488. DOI: 10.1093/nar/gkaa953.
Jiang, Y., Song, H., Gao, H., Zhang, Q., & Ding, X. (2022). Exploring the optimal strategy of imputation from SNP array to whole-genome sequencing data in farm animals. Frontiers in Genetics, 13, 963654. DOI: 10.3389/fgene.2022.963654.
All products and services are For Research Use Only and not for diagnostic or therapeutic use.
For research purposes only, not intended for clinical diagnosis, treatment, or individual health assessments.
Send a MessageFor any general inquiries, please fill out the form below.