When to Build a Population-Specific Imputation Reference Panel: Design, Sample Selection, and Validation

In modern population genomics, evolutionary biology, and complex disease genetics, genotype imputation is a widely used computational approach for increasing variant density in SNP microarray or low-pass whole-genome sequencing (lpWGS) datasets. Large public reference resources such as TOPMed and the Haplotype Reference Consortium (HRC) provide broad haplotype coverage, but global reference resources remain uneven in their representation of many regional and underrepresented populations. When the target cohort is genetically distant from the populations represented in a public panel, imputation accuracy can decline, particularly for low-frequency and rare variants (MAF < 5%).
For population-genetics consortia, regional biobanks, and project managers, this raises a fundamental strategic question: When does a study justify building a dedicated, population-specific imputation reference panel, and how should it be engineered?
Recent empirical studies show that genetic proximity can outweigh sheer reference sample size for some underrepresented populations, especially when the study focuses on low-frequency and rare variants. In appropriate cohorts, a smaller population-matched reference panel can outperform a substantially larger but less closely matched public panel for selected frequency ranges and downstream analyses. This guide details the design, sample selection, sequencing depth optimization, phasing, and validation framework used to evaluate and construct a population-specific imputation reference panel.
TL;DR
- Population Match Can Outweigh Sheer Size: A population-matched reference panel can outperform a substantially larger public panel for selected low-frequency and rare variants when the target cohort is underrepresented in the public reference.
- Recognize When Public Panels Underperform: Consider custom panel construction when public-panel imputation shows persistently weak quality metrics, poor empirical concordance, or marked population-specific allele-frequency and haplotype mismatches. Thresholds should be defined for the study rather than applied universally.
- Optimize Sample Selection via PCA: Sample across the breadth of population structure, representing relevant geographic and genetic subgroups while managing cryptic relatedness according to a pre-defined kinship threshold.
- Balance High and Low Depth WGS: Mixed-depth designs can combine a high-depth core with additional lower-depth genomes to expand haplotype representation within a fixed sequencing budget; the optimal allocation is cohort- and project-dependent.
- Enforce Modern Phasing Standards: Assemble panels using advanced phasing algorithms like SHAPEIT5 or Beagle 5.4, validating panel accuracy with empirical leave-out WGS benchmarking (r2, IQS, and NRD).
The Reference Panel Dilemma: Mega-Scale Public Panels vs Custom Population-Specific Panels
Genotype imputation operates on the principle of haplotype matching: unobserved genotypes in a target study sample are statistically inferred by identifying shared or closely matching haplotypes represented within a densely characterized reference panel.
1.1 Why Public Panels Can Underperform in Underrepresented Populations
When target cohorts originate from East Asian, Southeast Asian, African, Indigenous American, Middle Eastern, or Oceanian populations, public panels present three major vulnerabilities:
- Haplotype Sparsity in Distant Populations: Even very large public panels may contain limited representation of specific regional populations. When closely related haplotypes are sparse, imputation of population-specific low-frequency and rare alleles can be less accurate (Cengnata et al., 2024).
- Missing Population-Specific Variants: Population-specific alleles that are absent from a reference panel cannot be recovered from that panel through standard genotype imputation. This limitation becomes increasingly important as analyses move toward lower-frequency variants.
- Quality-Metric Calibration in Multi-Ancestry Panels: Estimated imputation-quality metrics can be upwardly biased relative to empirical dosage accuracy for ancestry groups that are less well represented in a multi-ancestry panel. This makes empirical validation especially important when relying on marginal-quality variants (Shi et al., 2024).
1.2 The "Build vs. Borrow" Decision Matrix
Project managers and consortia should evaluate the following objective criteria when deciding whether to construct a dedicated reference panel:
| Decision Criterion | Use Public Reference Panel (e.g., TOPMed / 1KGP) | Build Custom Population-Specific Panel |
| Cohort Ancestry | Cohorts that are well represented by an available public reference panel. | Underrepresented regional populations, indigenous cohorts, isolated founder groups, or non-human agricultural species. |
| Target Variant Frequency | Focus is exclusively on common variants (MAF ≥ 5%) for general GWAS or baseline PRS. | Focus includes low-frequency (1% ≤ MAF < 5%) and rare variants (MAF < 1%), fine-mapping, or loss-of-function discovery. |
| Preliminary Imputation Metrics | Empirical and estimated imputation metrics meet project-defined acceptance criteria across the frequency ranges of interest. | Empirical concordance is materially lower than expected, quality drops across priority loci, or the public panel poorly represents the target population. |
| Target Cohort Size | Small discovery projects (N < 1,000 samples) with constrained budgets. | Large or longitudinal cohorts where improvements in imputation quality can affect many downstream analyses. |
| Genotyping Scaffold | Standard high-density commercial arrays (e.g., Global Screening Array >700k SNPs). | Low-density custom arrays, reduced-representation sequencing, or ultra-low-pass WGS (0.5×–1×). |
When evaluating sequencing scaffolds or assessing baseline array technologies, review our guide on SNP Arrays vs Low-Pass and Deep WGS in Population Genomics.
Figure 2. Strategic PCA-guided donor selection capturing maximal ancestral diversity across regional populations.
Cohort Sampling and Genetic Diversity: Selecting the Ideal Reference Anchor Individuals
The ultimate accuracy of a custom reference panel is governed by the genetic diversity and representativeness of its constituent samples. Selecting biased, geographically clustered, or cryptically related individuals severely degrades the panel's generalizability across the broader target population.
2.1 PCA-Guided Stratified Sampling
Rather than randomly selecting available biobank samples, project leaders must apply stratified sampling across genetic principal component (PC) space:
- Capture the Breadth of PC Space: Select individuals that represent the major clusters, gradients, and less common regions of the population's PC space. Sampling should reflect the genetic and geographic structure relevant to the study rather than assuming that social, linguistic, or administrative categories map directly to genetic groups.
- Avoid Geographic or Sampling Imbalance: Biobanks can over-represent readily accessible recruitment sites. Where scientifically justified and ethically appropriate, donor selection should cover regional groups and genetic substructure that are relevant to the target cohort (Yang et al., 2025).
- Leverage Multi-Dimensional Clustering: Integrating Population Structure Analysis Service and PCA QC for GWAS: Outlier & Stratification Detection Guide can help identify unintended ancestry imbalance, outliers, and population-structure mismatch relative to the intended panel design.
2.2 Relatedness Filtering and Family Trios
- Cryptic Relatedness Pruning: For standard population reference panels, identify related sample pairs with tools such as KING or PLINK 2.0 and exclude or down-weight relatives according to the project design. In KING-style interpretation, kinship values below approximately 0.044 are commonly treated as more distant than third-degree relationships, while higher values indicate progressively closer relationships.
- Strategic Inclusion of Parent-Offspring Trios (When Feasible): While unrelated individuals maximize population representation, a subset of parent-offspring trios can provide transmission-informed phasing and an empirical benchmark for switch-error assessment. The number of trios should be selected according to panel size, sample availability, and validation goals.
2.3 Reference Donor Selection Checklist
Before committing biospecimens to deep whole-genome sequencing, every candidate donor must pass the following verification criteria:
| Verification Step | Quality Threshold | Purpose & Risk Mitigation |
| Ancestry Verification | Project-defined population-representation criterion based on PCA, ancestry modeling, sampling geography, and study objectives | Reduces unintended ancestry imbalance and ensures the panel reflects the population structure relevant to the target cohort. |
| DNA Integrity (DIN) | High-quality DNA suitable for the selected WGS library workflow; project-specific integrity criteria | Supports consistent library construction while recognizing that acceptable DNA integrity depends on the sequencing and library workflow. |
| Kinship Coefficient (φ) | Project-defined kinship threshold; values below approximately φ = 0.044 are commonly treated as more distant than third-degree relationships in KING-style interpretation | Limits redundant sampling when the goal is broad population haplotype representation, while allowing deliberate pedigree inclusion for phasing validation. |
| Population Representation | Balanced coverage of the geographic and genetic structure relevant to the target cohort | Reduces ascertainment bias and improves transferability across the intended study population. |
Sequencing Strategy and Depth Allocation: High-Depth vs. Low-Depth Blends
A common misconception in reference panel construction is that every donor must be sequenced at the same very high depth. Empirical work in Japanese population reference panels shows that mixed-depth designs can expand haplotype representation efficiently within a fixed sequencing budget, although the optimal balance depends on cohort composition and analysis goals (Flanagan et al., 2024).
3.1 The Power of Mixed-Depth Reference Architectures
In the Biobank Japan benchmark (Flanagan et al., 2024), researchers evaluated four population-specific reference panel configurations against TOPMed:
- 1KG augmented with 1,037 Japanese samples at 30×: Improved population representation relative to 1KGP alone, although performance gains varied across allele-frequency bins.
- 1KG augmented with 3,256 Japanese samples at 15–30×: Improved imputation performance for low-frequency variants in the Japanese target population compared with less closely matched public resources.
- 1KG augmented with 4,216 Japanese samples at approximately 3×: Showed that a larger low-pass contribution can capture useful population-specific haplotypes when supported by appropriate variant calling and phasing.
- 1KG augmented with 7,472 Japanese samples across mixed depths: Produced the strongest overall performance among the evaluated population-specific configurations, with reported gains over TOPMed for selected imputation metrics and downstream association analyses.
Key Architectural Takeaway: When project funds are constrained, a mixed-depth design that combines a high-depth core with a larger set of medium- or low-depth genomes can be evaluated as an alternative to sequencing every donor deeply. The optimal allocation should be tested for the target population rather than copied directly from another cohort. For sequencing support, see Whole Genome Re-sequencing for Population Genetics.
3.2 Library Chemistry: PCR-Free Illumina vs Long-Read Supplements
- PCR-Free Short-Read WGS: PCR-free library preparation reduces PCR-associated amplification bias and can improve coverage uniformity, although GC content, DNA quality, library construction, and sequencing platform can still influence difficult regions.
- Long-Read Integration: Long-read sequencing can improve discovery and haplotype representation of structural variants, segmental duplications, and difficult repeat regions. If long-read genomes are incorporated into a reference-panel strategy, any downstream gain in structural-variant imputation should be evaluated empirically for the target cohort.
For high-level institutional infrastructure and long-term cohort scaling, see our overview on Biobank Sequencing Strategy: From Array to WGS.
Figure 3. Conceptual resource-allocation comparison between uniform high-depth sequencing and mixed-depth population reference-panel strategies.
High-Accuracy Phasing and Panel Assembly Workflows
Once high-quality variant calls (SNPs and short InDels) are generated and filtered, the unphased diploid genotypes must be computationally phased into discrete, chromosome-length maternal and paternal haplotypes. Phasing errors directly introduce chimeric switch errors that degrade downstream imputation accuracy.
4.1 Modern Phasing Algorithms: SHAPEIT5 vs Beagle 5.4
- SHAPEIT5: A phasing framework designed for large sequencing datasets (Hofmeister et al., 2023). It uses separate workflows for common and rare variants and was shown to phase very rare variants with low switch-error rates in large UK Biobank sequencing datasets; observed error rates depend strongly on allele frequency, sample size, and data quality.
- Beagle 5.4: Highly efficient hidden Markov model phasing that natively handles unphased reference VCFs, providing rapid multi-threaded processing.
- Genetic Recombination Maps: Use an appropriate recombination or genetic map when required by the selected phasing workflow, ideally one that is suitable for the target population and genome build. Linkage-disequilibrium patterns can provide complementary population context through our Linkage Disequilibrium Analysis Service, but an LD analysis is not itself a substitute for a validated meiotic recombination map.
4.2 Handling Complex Genomic Loci
Specialized chromosomal regions require dedicated pre-processing:
- The MHC / HLA Region (Chr 6p21): The human leukocyte antigen locus harbors extreme polymorphism and long-range linkage disequilibrium. Phase the MHC locus as an isolated high-density chunk with customized genetic distance parameters.
- Sex Chromosomes (Chr X and Y): Males must be phased as haploid in the non-pseudoautosomal regions (non-PAR) and females as diploid. Pseudoautosomal regions (PAR1 and PAR2) must be split and processed as autosomes.
Rigorous Panel Validation & Empirical Benchmarking
A newly assembled reference panel must be empirically validated before release to ensure it delivers genuine accuracy gains over existing public resources.
5.1 Validation Methodologies
- Masked WGS Leave-Out Validation: Set aside a representative high-depth WGS test subset that is not included in the reference panel. Mask or down-sample markers to emulate the intended target dataset, impute with the custom and comparator panels, and compare imputed dosages against observed WGS genotypes across frequency bins and genomic regions. The number of held-out samples should be selected according to cohort size and validation objectives (Mauleekoonphairoj et al., 2023).
- Key Quantitative Validation Metrics:
- Concordance r2: The squared Pearson correlation between imputed dosage and true diploid genotype (0, 1, 2).
- Imputation Quality Score (IQS): An accuracy metric based on true/false positive/negative rates that corrects for chance agreement on rare alleles.
- Non-Reference Discordance Rate (NRD): The proportion of discordant non-reference genotypes or alleles, interpreted by allele-frequency bin and the validation design rather than against a single universal threshold.
- Rare Variant Performance: Evaluate empirical dosage r2, concordance, and recovery across low-frequency and rare-variant bins using thresholds defined for the intended downstream analysis.
To downstream post-imputation data filtering protocols and dosage extraction standards, refer to our technical companion guide on Post-Imputation QC for Large Cohorts: INFO/R², MAF, Concordance, and Dosage Filtering.
Figure 4. Conceptual leave-out validation framework comparing empirical imputation performance across allele-frequency tiers.
End-to-End Roadmap and Production Release Checklist
Constructing an institutional or population-scale imputation reference panel requires coordinated sampling, sequencing, variant processing, phasing, validation, and data governance. A documented operational checklist improves traceability and downstream reuse.
| Phase & Milestone | Operational Deliverable & Quality Gate | Key Tools & Standards |
| 1. Donor Selection & QC | Representative donors selected across the relevant population PC space; DNA integrity and kinship criteria defined for the project. | KING, PLINK 2.0, Fluorometric Qubit / TapeStation |
| 2. High-Throughput Sequencing | High-depth or mixed-depth WGS strategy selected after pilot review; sequencing quality and duplicate-rate criteria defined for the chosen workflow. | Illumina NovaSeq X Plus / PacBio Revio; FastQC, Picard |
| 3. Joint Genotyping & Filtering | High-confidence, normalized variant set with multi-allelic handling and species-appropriate variant QC metrics. | BWA-MEM2, DeepVariant / GATK, bcftools |
| 4. High-Resolution Phasing | Chromosome-wide phased haplotypes with switch-error assessment in benchmark samples or pedigrees where available. | SHAPEIT5 (phase_common + phase_rare), Beagle 5.4 |
| 5. Leave-Out Empirical Audit | Held-out WGS validation sized to the cohort; empirical dosage r2, concordance, and NRD evaluated by allele-frequency bin. | Custom validation scripts, GLIMPSE2_concordance (Biagini et al., 2025) |
| 6. Multi-Format Release | Reference-panel and target-data formats compatible with the selected imputation software; indexed files and versioned metadata for reproducible deployment. | Minimac4, GLIMPSE2, IMPUTE5, tabix or other project-appropriate tools |
Where custom genotyping arrays are utilized in conjunction with imputation panels, cross-platform harmonization can be evaluated via SNP Genotyping Service workflows. For panel construction, deployment, phasing, and imputation support, see our Genotype Imputation & Haplotype Phasing Service. In downstream association studies, Genome-wide Association Analysis Service (GWAS) can support population-specific association testing. For complex livestock, crop, or polyploid species, exploring Combining Reduced Representation Genome Sequencing with Imputation: When It Works and How to Validate provides additional validation frameworks.
Planning note: Sample numbers, sequencing depths, kinship cutoffs, imputation-quality metrics, and validation thresholds in this article are examples for study planning rather than universal acceptance criteria. Appropriate values depend on species, population structure, reference-panel composition, sequencing technology, genome build, allele-frequency range, and downstream analysis objectives. Project-specific criteria should be defined and validated before full-scale panel deployment.
FAQs
There is no universal minimum. Population-specific panels ranging from hundreds to several thousand well-selected individuals have been useful in different settings, but the required sample size depends on population diversity, target variant frequencies, sequencing depth, and the size and composition of available comparator panels. Pilot validation against held-out samples is more informative than applying a fixed number across projects.
Imputation relies on haplotypes represented in the reference panel. Low-frequency and rare variants are often more geographically or population restricted than common variants, so a smaller but better-matched panel can provide more relevant haplotypes than a much larger panel with limited representation of the target population. Whether it outperforms the larger panel should be tested empirically.
Yes, low-pass genomes can contribute useful haplotype information when variant calling and phasing are designed for low-coverage data. Population-specific studies have shown that mixed-depth reference architectures can perform well, but the optimal balance of high- and low-depth samples should be evaluated for the target cohort rather than assumed in advance.
Close relatives can reduce the number of independent haplotypes represented when the panel is intended to maximize population diversity. Related sample pairs should be identified with a project-defined kinship rule; in KING-style interpretation, kinship below approximately 0.044 is commonly treated as more distant than third-degree relationships. Pedigree samples can still be retained deliberately when transmission-based phasing or validation is part of the design.
Use a validated recombination or genetic map that is compatible with the selected phasing workflow and genome build. A population-matched map can be useful when available, but the benefit should be assessed against established alternatives. Linkage-disequilibrium maps and LD analyses provide related population-genetic context but are not interchangeable with meiotic recombination maps.
A strong validation strategy is to hold out representative high-depth WGS samples, emulate the intended target-data density, impute those samples with the custom and comparator panels, and compare imputed dosages with observed WGS genotypes. Report empirical dosage r2, concordance, and non-reference discordance across allele-frequency bins; the size of the held-out set should reflect the cohort and validation objective.
Potentially. Population-specific haplotypes can be combined or used alongside public reference resources when file formats, genome builds, variant representation, consent or data-use conditions, and phasing strategy are compatible. The benefit of augmentation should be measured empirically because a larger merged panel does not automatically improve every frequency range or population.
Next steps: If you're planning a custom reference panel project or evaluating panel feasibility for an underrepresented cohort, explore our Genotype Imputation & Haplotype Phasing Service to review target data, reference-panel options, phasing strategy, and validation requirements.
References:
- Cengnata, A., et al. "A genotype imputation reference panel specific for native Southeast Asian populations." NPJ Genomic Medicine, 2024.
- Flanagan, J., et al. "Population-specific reference panel improves imputation quality for genome-wide association studies conducted on the Japanese population." Communications Biology, 2024.
- Mauleekoonphairoj, John, et al. "A diverse ancestrally-matched reference panel increases genotype imputation accuracy in a underrepresented population." Scientific Reports, 2023.
- Hofmeister, R. J., et al. "Accurate rare variant phasing of whole-genome and whole-exome sequencing data in the UK Biobank." Nature Genetics, 2023.
- Yang, Qingxin, et al. "High-quality Population-specific Haplotype-resolved Reference Panel in the Genomic and Pangenomic Eras." Genomics, Proteomics & Bioinformatics, 2025.
- Biagini, Simone Andrea, et al. "Genotype imputation from low-coverage data for medical and population genetic analyses." Genome Research, 2025.
- Shi, M., et al. "Genotype imputation accuracy and the quality metrics of the minor ancestry in multi-ancestry reference panels." Briefings in Bioinformatics, 2024.