Can an ~85K SNP Array Support Genotype Imputation? What to Check Before Genotyping
Figure 1. A usable imputation scaffold requires compatible markers, a suitable reference panel, and a harmonization plan before samples are genotyped.
Yes, an approximately 85K SNP array can support genotype imputation, but suitability depends on more than the nominal marker count. The observed markers must overlap the selected reference panel, tag haplotypes in the study population, span the genome without consequential gaps, and survive build, strand, allele, and genotype-quality checks. The downstream objective also matters because a scaffold that is adequate for common variants may perform poorly for low-frequency or population-specific variants.
The right time to test these conditions is before genotyping. A pre-project evaluation can compare the actual array manifest with candidate reference panels, identify ancestry or region-specific weaknesses, and define a small validation set. This protects the cohort from a common failure mode: technically clean array data that cannot deliver the imputed variant set the analysis requires.
TL;DR
- An 85K count does not predict imputation accuracy by itself.
- Reference-panel overlap and ancestry match are first-order requirements.
- Genome build, chromosome naming, strand, and allele coding must be harmonized before phasing.
- Marker spacing and allele frequency determine which haplotypes the scaffold can resolve.
- Validate with masked markers or sequenced samples and report performance by frequency and ancestry, not only as one average.
Start With the Target
Define what the imputed data must support. A common-variant GWAS, cross-cohort meta-analysis, fine mapping, HLA-focused work, and research on low-frequency population-specific variation have different accuracy and coverage needs. "More imputed variants" is not a sufficient endpoint because poorly imputed dosages can add noise without adding usable information.
Write a target specification before selecting the array:
- Variant classes and allele-frequency range needed for analysis
- Populations and ancestry groups represented in the cohort
- Genome build and coordinate convention used by downstream data
- Required chromosomes and any special regions
- Intended use of dosages, hard calls, or genotype probabilities
- Minimum acceptable concordance or imputation-quality distribution for the research endpoint
- Whether the data must merge with an existing genotyped cohort
The Genotype Imputation and Haplotype Phasing service can be scoped against this specification. Researchers who are still deciding among array and sequencing strategies can review SNP arrays versus low-pass and deep WGS.
The Workflow Has Six Gates
Imputation is often summarized as a software step. In practice, each upstream gate constrains what the software can infer.
| Gate | Main question | Evidence to retain |
| Array genotyping | Are sample and marker calls reliable? | Call-rate, cluster, heterozygosity, duplicate, and plate summaries |
| Pre-imputation QC | Are identifiers, coordinates, and alleles consistent? | Exclusion log and corrected manifest |
| Harmonization | Do build, strand, and allele representations match the reference? | Matched, flipped, swapped, ambiguous, and rejected marker counts |
| Phasing | Can study haplotypes be estimated from the observed scaffold? | Phasing logs, chromosome outputs, and sample order checks |
| Imputation | Does the reference panel supply compatible haplotypes? | Panel version, software version, parameters, and dosage outputs |
| Post-imputation QC | Which inferred variants are usable for the endpoint? | INFO or r² distributions, frequency bins, concordance, and filters |
Figure 2. Imputation quality is cumulative: a mismatch introduced before phasing propagates into every downstream dosage.
The site's QC metrics at cohort scale guide explains why sample, marker, plate, and batch summaries should be reviewed before harmonization. The VCF Generation service can also be relevant when the delivery format and allele representation need to integrate with a downstream workflow.
Reference-Panel Overlap Comes First
An array marker helps imputation only if it can be matched reliably to the reference panel and contributes information about neighboring haplotypes. The initial overlap calculation should use the exact manifest and the exact panel release, not a marketing marker count or a previous array version.
Count markers through a staged reconciliation:
- Map all array markers to the intended genome build.
- Match by chromosome and position, then confirm reference and alternate alleles.
- Identify strand flips, allele swaps, multi-allelic positions, and duplicate coordinates.
- Separate unambiguous matches from A/T and C/G markers that require frequency or strand evidence.
- Recalculate overlap after technical QC and after unresolved markers are removed.
The retained overlap should then be summarized by chromosome and genomic interval. A good whole-genome percentage can conceal long gaps or weak coverage in regions central to the research objective. For a general comparison of candidate panels, see choosing a reference panel for genotype imputation.
Ancestry Determines Haplotype Match
Reference size matters, but representation matters with it. A large panel may contain many haplotypes while still underrepresenting the ancestry or admixture pattern of the study cohort. Sengupta and colleagues reported marked differences in imputation outcomes among reference panels for sub-Saharan African data. Xu and colleagues showed that population-specific add-on markers and an internal panel could improve results in an underrepresented population.
For a multi-ancestry cohort, do not evaluate imputation only across all samples. Stratify validation by genetically informed ancestry group and, when relevant, examine admixed samples separately. The goal is not to assign rigid labels to individuals. It is to detect whether the observed scaffold and reference haplotypes perform unevenly across the cohort.
Useful pre-genotyping evidence includes:
- Principal-component comparison with suitable reference populations
- Expected marker polymorphism and MAF distributions by group
- Array-reference overlap by ancestry-relevant panel subset
- Known gaps in population representation
- A plan for mixed or uncertain ancestry rather than forced exclusion
The Population Structure Analysis Service can support this characterization, and the PCA QC guide for GWAS explains how structure and outliers affect downstream analysis.
Spacing Beats a Global Count
Imputation draws information from local haplotype structure. Consequently, the same number of markers can behave differently when distributed evenly, concentrated around selected loci, or separated by gaps in high-recombination regions. Marker density should be summarized along the genetic map as well as physical coordinates because a fixed number of base pairs does not represent the same recombination distance everywhere.
Review at least four spacing views:
| View | What it reveals | Decision use |
| Markers per chromosome | Gross imbalance or missing chromosomes | Detect configuration problems |
| Physical gap distribution | Long unobserved intervals | Find regions with limited scaffold support |
| Genetic-map spacing | Sparse coverage in high-recombination regions | Estimate haplotype resolution |
| LD with target variants | Population-specific tagging | Test whether key variants are inferable |
If the project prioritizes reported loci or poorly covered regions, a targeted supplementation review may be more efficient than replacing the whole backbone. Supplementation still requires probe feasibility and cannot be promised for every requested marker. The SNP Genotyping Service provides a broader route for evaluating alternative marker configurations.
Build and Strand Errors Are Preventable
Many imputation failures begin as metadata errors. Coordinates from different reference builds can appear plausible while mapping to different loci. A/T and C/G variants are strand-ambiguous without additional information. Reference and alternate allele conventions can change across files, and legacy identifiers may map to multiple or retired records.
A pre-imputation manifest should preserve chromosome, position, rsID when available, reference allele, alternate allele, reported array allele, strand convention, and build. Harmonization outputs should not merely state that markers were "aligned." They should show how many were matched directly, flipped, allele-swapped, lifted over, rejected as ambiguous, or removed as duplicates.
Stop and investigate when:
- Allele-frequency comparisons show a systematic inversion or large unexplained differences.
- One chromosome loses a disproportionate share of markers.
- Many positions require liftover but lack stable identifiers.
- Ambiguous SNPs cannot be resolved with trusted frequency information.
- Sample identifiers or chromosome encodings change between pipeline stages.
These checks can be planned alongside the Human 85K SNP Genotyping Array Service, so the delivered manifest and genotype files are compatible with the intended imputation route.
Candidate or Poor Candidate
Two cohorts genotyped at the same nominal density can produce different imputation outcomes because their ancestry, sample QC, marker overlap, and reference resources differ.
| Project profile | Assessment | Reason |
| Large cohort, common-variant GWAS, well-matched panel, strong overlap | Good candidate | The scaffold and reference jointly support common haplotypes |
| Multi-ancestry cohort with unequal group sizes | Needs evaluation | Performance may vary substantially among groups |
| Existing cohort must merge with a different array | Needs evaluation | Shared scaffold and harmonization may limit the joint variant set |
| Population poorly represented in public panels | Needs evaluation | Local haplotypes and low-frequency variants may be missed |
| Primary goal is novel rare-variant discovery | Poor candidate | Imputation predicts from reference haplotypes rather than discovering unobserved variants |
| Many key regions lack overlapping markers | Poor candidate without redesign | Local gaps cannot be repaired by a larger distant reference panel |
Figure 3. Cohort and reference-panel compatibility, not nominal array density, separates good candidates from projects that need redesign.
Validate Before Full Genotyping
A useful pilot includes representative samples from each ancestry group, study site, extraction batch, and input-quality range. If sequence truth data exist for some participants, compare imputed dosages or calls with those genotypes. Otherwise, mask a subset of directly genotyped markers, impute them, and compare the inferred values with the observed data.
Report validation by chromosome, ancestry group, and allele-frequency bin. The overall average can look strong while the variants most relevant to the project perform poorly. Inspect INFO or r² distributions, non-reference concordance, allele-frequency differences, and the number of variants remaining after post-imputation thresholds. The post-imputation QC guide for large cohorts covers these downstream filters, while why genotype imputation fails is intended for projects that already have problematic outputs.
The project is ready to scale when the pilot answers four questions: Does harmonization retain enough informative markers? Is performance acceptable in every important cohort group? Are the target allele-frequency ranges supported? Can the pipeline be reproduced with versioned inputs and documented exclusions?
Prepare a Review Package
A feasibility review is faster when the technical inputs arrive as one versioned package. Include the array manifest rather than a product nickname, because platform names can cover multiple releases or custom configurations. State whether coordinates have already been converted between builds and retain the source build even when a lifted file is supplied.
The review package should contain the proposed sample count, ancestry or population description, primary downstream analysis, and the allele-frequency range that matters. Add the candidate reference panels and any access constraints. If the cohort must merge with existing data, include the existing platform, build, sample count, and a marker-level overlap summary. When a sequenced subset or orthogonal genotype set exists, describe how those samples map to the new cohort and whether they represent all major groups.
Record estimates before genotyping by labeling them clearly. Expected overlap based on a manifest is not the same as retained overlap after actual QC. A projected imputation yield is not a validated accuracy result. The final project plan should therefore separate three milestones: design-time compatibility, pilot performance, and production performance.
Request these outputs from the pre-project review:
- A manifest-to-panel overlap table by chromosome
- A list of unresolved build, strand, allele, and identifier issues
- Marker-spacing summaries and regions requiring attention
- An ancestry-aware pilot sampling plan
- Masking or sequence-truth validation methods
- Planned imputation-quality summaries by MAF and cohort group
- Clear go, revise, or select-another-platform criteria
This package turns an abstract question about 85K markers into a reproducible decision that can be revisited if the array version, cohort composition, or reference panel changes.
Troubleshoot the Pilot in a Fixed Order
Suppose a pilot produces a reasonable genome-wide INFO distribution but weak results on one chromosome and in one ancestry stratum. Start with inputs before changing algorithms. Confirm sample identities, the manifest release, genome build, chromosome naming, reference and alternate alleles, strand treatment, duplicate variants, and reference-panel overlap. Next, inspect observed-marker missingness and spacing in the affected region. Only then compare phasing or imputation settings. This sequence helps distinguish a coordinate or allele error from a biological coverage limitation and prevents parameter tuning from hiding an upstream defect.
Use controls that answer different questions. Masked array markers test recovery at loci the platform can already observe; a sequenced subset can test variants that were never on the array; replicate samples can reveal workflow reproducibility. None is a complete substitute for the others. Report results by ancestry, chromosome, allele-frequency range, and target use case, and preserve denominators so that filtering does not make a weak pilot appear stronger. Published quality cutoffs are useful reference points, but the project should justify its own acceptance rules from the planned analysis and the consequences of genotype uncertainty.
If performance remains weak after harmonization, ask whether the failure is local or global. A local gap may support custom supplementation or direct sequencing of a priority region. Broad ancestry mismatch or sparse coverage may favor another genome-wide platform. Imputation is useful when a compatible observed scaffold and reference panel support the variants of interest; it adds little when the target variants are absent from or poorly represented in the reference. Imputed association is still association, not proof of causality, and important findings may require independent replication or direct genotyping.
Frequently Asked Questions
There is no universal marker-count threshold. Accuracy depends on marker distribution, LD, ancestry, array-reference overlap, genotype quality, and the frequency of the variants being imputed. A pilot using the actual manifest and panel is more informative than the count.
Choose by empirical performance in a population similar to the study cohort, not only by panel size. Consider genome build, represented ancestries, variant classes, access conditions, and whether the panel supports the required chromosomes and allele-frequency range.
They can sometimes be harmonized and imputed to a shared reference, but the joint analysis is constrained by compatible markers, batch effects, and platform-specific missingness. Evaluate each array separately before merging dosages or association results.
A/T and C/G SNPs require strand evidence, allele-frequency comparison, or validated manifests. Remove them when orientation cannot be resolved confidently. Keeping uncertain markers can introduce systematic allele errors.
Provide the exact array manifest, cohort size, ancestry composition, genome build, existing truth or sequence data, target analyses, candidate reference panels, and any loci that must be represented. Include information about sample source, extraction batches, and planned plate structure.
References
- Berdnikova AA, Zorkoltseva IV, Tsepilov YA, et al. Genotype imputation in human genomic studies. Vavilov Journal of Genetics and Breeding. 2024;28(6):628-639. doi:10.18699/vjgb-24-70.
- Sengupta D, Botha G, Meintjes A, et al. Performance and accuracy evaluation of reference panels for genotype imputation in sub-Saharan African populations. Cell Genomics. 2023;3(6):100332. doi:10.1016/j.xgen.2023.100332.
- Xu ZM, Rüeger S, Zwyer M, et al. Using population-specific add-on polymorphisms to improve genotype imputation in underrepresented populations. PLOS Computational Biology. 2022;18(1):e1009628. doi:10.1371/journal.pcbi.1009628.
- Byrska-Bishop M, Evani US, Zhao X, et al. High-coverage whole-genome sequencing of the expanded 1000 Genomes Project cohort including 602 trios. Cell. 2022;185(18):3426-3440.e19. doi:10.1016/j.cell.2022.08.004.
- Taliun D, Harris DN, Kessler MD, et al. Sequencing of 53,831 diverse genomes from the NHLBI TOPMed Program. Nature. 2021;590(7845):290-299. doi:10.1038/s41586-021-03205-y.
- Yang H-C, Kwok P-Y, Li L-H, et al. The Taiwan Precision Medicine Initiative provides a cohort for large-scale studies. Nature. 2025;648(8092):117-127. doi:10.1038/s41586-025-09680-x.
- Naj AC. Genotype Imputation in Genome-Wide Association Studies. Current Protocols in Human Genetics. 2019;102(1):e84. doi:10.1002/cphg.84.
For research purposes only. CD Genomics provides population-genomics genotyping and bioinformatics support for research, not for clinical diagnosis or individual medical decision-making.