Is an ~85K SNP Array Dense Enough for Your Human GWAS? A Pre-Project Evaluation Guide
Figure 1. GWAS suitability depends on effective genomic coverage, cohort composition, power, and the downstream analysis plan rather than marker count alone.
An approximately 85K SNP array can support some human genome-wide association studies, but 85,000 is not a universal adequacy threshold. The useful question is whether the assayed markers capture enough of the common haplotype structure in the study population, at the allele frequencies relevant to the phenotype, with a cohort large enough to detect the expected effects. The answer also depends on whether the study will use genotype imputation and whether the objective is broad discovery, replication, cohort characterization, or analysis of predefined loci.
For a relatively homogeneous cohort studying common variants, with an ancestry-matched reference panel and a validated imputation plan, an approximately 85K backbone may be a practical starting point. A diverse or admixed cohort, a rare-variant objective, a small expected effect, weak reference-panel coverage, or a need for direct novel-variant discovery usually argues for additional markers or a sequencing strategy. Teams can use the checkpoints below before requesting array processing or committing the whole cohort.
TL;DR
- Marker count is only a coarse label. Linkage disequilibrium, genomic spacing, allele frequency, ancestry, and probe performance determine effective coverage.
- Cohort size and phenotype quality often influence GWAS power more than a modest difference in array density.
- Imputation can expand the analyzable variant set, but only when the scaffold and reference panel align.
- Rare-variant discovery, structural variation, and poorly represented ancestries require a different evaluation.
- A representative pilot is the safest way to test call quality, population structure, and imputation performance before scale-up.
Marker Count Is Not Coverage
Two arrays with similar marker counts can deliver very different value for the same cohort. A marker contributes to GWAS coverage when it is informative in the study population, passes technical QC, and tags nearby variation through linkage disequilibrium. Markers that cluster in a few genomic regions, have very low frequency in the target population, or fail calling contribute less to genome-wide association power than their raw count implies.
The distinction is especially important at approximately 85K markers. Average spacing obtained by dividing genome length by marker count hides long gaps, uneven recombination, and local variation in LD. A better pre-project review uses the actual manifest, maps markers to the intended genome build, and examines spacing and tagging by chromosome and genomic region. Researchers who need a formal LD review can pair the array assessment with a linkage disequilibrium analysis service.
Effective coverage has several layers:
- Technical coverage. How many planned markers generate reliable genotype clusters and acceptable call rates in the submitted samples?
- Population coverage. Are those markers polymorphic and informative in the ancestries represented by the cohort?
- Haplotype coverage. Do assayed markers tag untyped variants through the cohort's LD structure?
- Analytical coverage. Does the retained marker set support the intended association model, covariates, and imputation workflow?
- Biological coverage. Are important loci, genomic regions, and allele-frequency ranges represented for the research question?
This is why the broader SNP arrays versus low-pass and deep WGS comparison should be read as a method-selection guide, not as a ranking based on the number of measured variants.
Study Goals Set the Bar
An array can be adequate for one objective and unsuitable for another. Define the primary endpoint before comparing platforms. Exploratory GWAS, replication of reported signals, ancestry-aware cohort QC, and rare-variant discovery impose different demands on the marker backbone.
| Study goal | What the array must provide | When ~85K may fit | Main reason to reconsider |
| Common-variant discovery GWAS | Broad tagging of common haplotypes, stable QC, adequate sample size | When population coverage and imputation are validated | Gaps in tagging, weak imputation, or small effects relative to cohort size |
| Replication or follow-up | Direct coverage or strong proxies for known loci | When target loci and proxies are on the manifest | Key variants or ancestry-specific proxies are absent |
| Population structure and relatedness | Genome-wide, LD-pruned informative markers | Often feasible after cohort-specific QC | Markers are monomorphic or unevenly informative across groups |
| Fine mapping | Dense local coverage around association signals | Only as a screening scaffold | Local resolution is too low for credible-set refinement |
| Rare-variant discovery | Direct observation of low-frequency and rare alleles | Generally not the primary fit | Fixed arrays do not discover most unassayed rare variants |
| Structural-variant discovery | Sequence or validated structural probes | Only for predefined array content | The study needs open-ended CNV or SV discovery |
The Genome-wide Association Analysis Service can be scoped after this endpoint is fixed. Association testing cannot rescue a backbone that misses the relevant variation, but a denser platform cannot rescue an underpowered cohort or poorly defined phenotype either.
Figure 2. Seven linked design factors determine whether an approximately 85K scaffold is informative for a particular cohort.
Ancestry Changes Effective Density
Array density is population-dependent because allele frequencies and LD patterns differ across ancestries. A marker selected as an efficient tag in one population may be rare, monomorphic, or weakly correlated with nearby variants in another. Mixed-ancestry and admixed cohorts add local ancestry variation, so a single whole-cohort summary can hide poor coverage in particular subgroups.
Recent reference resources illustrate why representation matters. The expanded high-coverage 1000 Genomes resource and TOPMed provide broad catalogs, while population-specific evaluations show that the largest panel is not automatically the most accurate for every group. Sengupta and colleagues found material differences among reference panels across sub-Saharan African populations, and Xu and colleagues showed that population-specific add-on markers can improve imputation in an underrepresented cohort. These findings support a practical rule: evaluate the study cohort and the reference resource together.
Before approving the array, ask for three ancestry checks:
- Compare the manifest with population-specific allele frequencies and LD where suitable reference data exist.
- Estimate how many markers remain informative after expected QC and LD pruning within each major ancestry group.
- Plan population-structure analysis before association testing, including how principal components, relatedness, and outliers will be handled.
The site's PCA analysis service and population structure analysis service provide natural downstream routes when cohort heterogeneity is a central design issue. They should be planned before genotyping rather than added only after an unexpected cluster appears.
Power Depends on More Than Density
GWAS power is a joint function of cohort size, allele frequency, effect size, phenotype measurement, study design, missingness, and the correlation between an assayed or imputed marker and the causal variant. Increasing marker density may improve tagging, but it does not create more informative participants or repair noisy phenotypes. Uffelmann and colleagues describe GWAS as an integrated sequence of design, QC, imputation, association, and interpretation decisions rather than a single laboratory measurement.
A pre-project power analysis should therefore vary more than one input. Use plausible effect sizes and allele frequencies rather than optimistic values. Evaluate the loss of usable samples after exclusions, the number of ancestry strata, relatedness, covariate burden, and any imbalance between comparison groups. For binary phenotypes, the smaller group often limits power. For quantitative traits, measurement reliability and distribution may dominate.
Three conclusions usually follow:
- A large, well-characterized cohort may obtain useful common-variant power from an 85K scaffold if effective coverage and imputation are adequate.
- A small cohort does not become well powered merely because the array is denser.
- When expected effects are small or the relevant variants are uncommon, both sample size and platform strategy may need to change.
The more general population genomics study-design guide is useful for separating sampling and sequencing decisions, while this page keeps the decision specific to an approximately 85K human array.
Imputation Extends the Scaffold
Imputation predicts untyped genotypes from observed array markers and reference haplotypes. It can increase the number of variants available for association and support meta-analysis across cohorts genotyped on different platforms. It does not turn every low-density array into an equivalent copy of high-coverage sequencing.
The scaffold must survive several gates: marker call quality, genome-build alignment, allele and strand harmonization, overlap with the chosen reference panel, phasing quality, and adequate LD between observed and imputed variants. Accuracy generally falls for lower-frequency variants and for populations that are poorly represented in the reference panel. Post-imputation filtering also reduces the final analyzable set.
| Imputation question | Evidence to review before scale-up | Warning sign |
| Does the panel overlap the array? | Exact marker intersection after build and allele alignment | Low overlap or many ambiguous variants |
| Does the panel match ancestry? | Validation by ancestry group and allele-frequency bin | One subgroup performs materially worse |
| Is marker spacing usable? | Chromosome-level gap and LD analysis | Long gaps in regions important to the study |
| Can harmonization be completed? | Build, chromosome, position, reference/alternate allele, strand | Many unresolved strand or identifier conflicts |
| Is accuracy adequate for the endpoint? | Masked-marker concordance and r²/INFO distributions | Acceptable average but poor low-frequency performance |
Researchers can review reference-panel selection for genotype imputation and the Genotype Imputation and Haplotype Phasing service. The separate article on using an approximately 85K scaffold for imputation focuses on these checks in greater depth.
Four Project Scenarios
Scenario testing prevents a generic yes-or-no answer from being applied to the wrong study. The examples below are decision patterns, not performance guarantees.
| Scenario | Likely assessment | What to validate first |
| Large, relatively homogeneous cohort; common quantitative trait; matched reference panel | Potentially suitable | Effective marker coverage, pilot imputation by MAF, and power after QC exclusions |
| Multi-ancestry cohort; uneven subgroup sizes; common trait | Needs evaluation | Per-group informativeness, local and global ancestry, reference-panel match, and stratified performance |
| Replication of reported loci in thousands of samples | Often suitable if loci are covered | Direct presence of each target SNP or a strong ancestry-appropriate proxy |
| Rare-variant or novel-variant discovery | Poor primary fit | Whether sequencing or targeted resequencing better addresses direct observation |
| Fine mapping around a known locus | Scaffold only | Local marker density, credible-set requirements, and targeted follow-up design |
Figure 3. Scenario-based evaluation turns a marker-count question into an evidence-based platform decision.
If the question centers on a defined region or a compact set of known variants, the Targeted Resequencing Service may be more direct. If the study requires broad open-ended variant discovery, whole-genome resequencing may provide a better data model despite its greater analytical burden.
Pre-Project Readiness Checklist
An array decision is ready for approval when the team can answer the following questions with evidence rather than assumptions:
- Cohort definition: Are inclusion rules, group labels, ancestry composition, relatedness policy, and expected final sample count documented?
- Phenotype definition: Are primary endpoints, covariates, missing-data rules, and coding conventions fixed before plate assignment?
- Manifest review: Has the exact marker manifest been mapped to the correct genome build and assessed for genomic spacing?
- Population relevance: Are allele frequencies and LD patterns appropriate for each represented population?
- Power model: Does the calculation use realistic effect sizes, allele frequencies, group imbalance, and post-QC sample counts?
- Imputation plan: Is there a named reference panel, harmonization route, and pilot validation design?
- Alternative strategy: Is there a clear trigger for custom supplementation, targeted resequencing, low-pass WGS, or standard-depth WGS?
- Deliverables: Are platform-native data, PLINK or VCF outputs, QC tables, and downstream analysis requirements specified?
The Human 85K SNP Genotyping Array Service supports project-specific review of the cohort, marker configuration, plate strategy, QC, and analysis-ready deliverables.
A Defensible Go Decision
"Dense enough" should mean that the selected platform meets a documented scientific and analytical requirement. It should not mean that the marker count looks large in isolation. A defensible go decision links the actual manifest to the cohort's ancestry and LD, demonstrates sufficient power for the planned endpoint, and confirms that imputation or follow-up sequencing can address known gaps.
For uncertain projects, run a representative pilot that includes the major ancestry groups, sample sources, extraction batches, and study groups. Review call quality, missingness, heterozygosity, population structure, relatedness, marker spacing, and masked-marker imputation performance. Scale only after the pilot produces evidence that the full cohort will answer the intended research question.
What the Array Result Can and Cannot Establish
Consider a research cohort assembled from two ancestry groups, with one group contributing fewer participants and showing shorter LD around several priority loci. A pooled call-rate summary could look acceptable while the smaller group has weaker local coverage or imputation. The practical response is not to apply a published sample size or imputation threshold as a universal rule. Evaluate retained markers, allele-frequency spectra, spacing, and masked-marker performance within each group, then decide whether the main analysis should be stratified, meta-analyzed, supplemented, or moved to a sequencing design. Record that decision before examining association results so that later choices are auditable.
A statistically associated array marker identifies a locus for further investigation; it does not by itself identify a causal variant or biological mechanism. Correlated variants, population structure, phenotype error, batch imbalance, and chance can produce or distort signals. Replication in an independent research dataset, conditional or fine-mapping analysis, and functional work may be required, depending on the claim. An 85K array may add little value when the primary objective is unbiased rare-variant discovery, when priority regions have sparse informative markers in the cohort, or when no suitable imputation resource exists. In those cases, changing the platform is a design correction rather than an array failure.
Frequently Asked Questions
It can be enough for some common-variant GWAS designs, especially in large cohorts with suitable LD coverage and validated imputation. The answer changes with ancestry, allele frequency, effect size, phenotype quality, and the final marker set after QC.
No. Imputation depends on the observed scaffold, reference-panel overlap, ancestry match, harmonization, and haplotype structure. It can extend useful coverage, but poor scaffold design or a mismatched reference panel will remain limiting.
Not automatically. A diverse cohort needs per-group evaluation of marker informativeness and imputation performance. An array, a supplemented array, low-pass WGS, or deeper sequencing may each be reasonable depending on the objective and available reference resources.
A fixed array directly observes only predefined loci. Rare variants absent from the manifest are not discovered, and imputation accuracy usually declines as frequency falls. Direct sequencing is preferable when rare-variant discovery is central.
Include samples representing the main study groups, ancestries, sites, extraction batches, and input-quality range. The pilot should test laboratory QC, population structure, harmonization, and imputation feasibility against prespecified acceptance criteria.
References
- Uffelmann E, Huang QQ, Munung NS, et al. Genome-wide association studies. Nature Reviews Methods Primers. 2021;1(1):59. doi:10.1038/s43586-021-00056-9.
- Yang H-C, Kwok P-Y, Li L-H, et al. The Taiwan Precision Medicine Initiative provides a cohort for large-scale studies. Nature. 2025;648(8092):117-127. doi:10.1038/s41586-025-09680-x.
- Sharma S, Nagar SD, Pemu P, et al. Genetic ancestry and population structure in the All of Us Research Program cohort. Nature Communications. 2025;16(1):4123. doi:10.1038/s41467-025-59351-8.
- Byrska-Bishop M, Evani US, Zhao X, et al. High-coverage whole-genome sequencing of the expanded 1000 Genomes Project cohort including 602 trios. Cell. 2022;185(18):3426-3440.e19. doi:10.1016/j.cell.2022.08.004.
- Sengupta D, Botha G, Meintjes A, et al. Performance and accuracy evaluation of reference panels for genotype imputation in sub-Saharan African populations. Cell Genomics. 2023;3(6):100332. doi:10.1016/j.xgen.2023.100332.
- Xu ZM, Rüeger S, Zwyer M, et al. Using population-specific add-on polymorphisms to improve genotype imputation in underrepresented populations. PLOS Computational Biology. 2022;18(1):e1009628. doi:10.1371/journal.pcbi.1009628.
- Truong VQ, Woerner JA, Cherlin TA, et al. Quality Control Procedures for Genome-Wide Association Studies. Current Protocols. 2022;2(11):e603. doi:10.1002/cpz1.603.
For research purposes only. The content and services described here are not intended for clinical diagnosis, treatment decisions, or individual health assessment.