Human SNP Array Project Design: Cohort Size, Plate Layout, and Batch Strategy Before Sample Submission
Figure 1. Array project design should distribute biological groups across technical factors before samples enter the laboratory.
The most expensive SNP-array batch effect is the one built into the plate map. If cases, controls, ancestry groups, sites, or extraction batches are separated across plates, a technical difference can become indistinguishable from the biological contrast. Statistical adjustment may reduce some artifacts, but it cannot reliably recover information when the factor of interest is perfectly confounded with processing.
A defensible design starts with the cohort and metadata, then creates the plate layout. Determine the primary analysis, expected final sample count, group balance, ancestry structure, relatedness policy, extraction history, and control strategy before sample submission. The goal is not randomization for its own sake. It is to make every important biological comparison estimable without relying on assumptions about laboratory batches.
TL;DR
- Size the cohort for the analysis after expected QC exclusions, not for the number of samples collected.
- Balance study groups, ancestry strata, sites, extraction batches, and major covariates across plates wherever practical.
- Randomize sample positions within those constraints and preserve the seed or assignment record.
- Use bridge samples and selected technical replicates to measure consistency across plates or processing waves.
- Freeze a versioned manifest that links sample IDs, metadata, plate positions, and downstream analysis files.
Define the Analysis Population
Begin with the participants or research samples that must contribute to the primary analysis. A collection database often contains duplicates, relatives, pilot samples, missing phenotypes, consent restrictions, or samples with insufficient DNA. The genotyping count should distinguish the number available, the number submitted, and the number expected to remain after laboratory and analytical QC.
For an association study, document:
- Primary phenotype and comparison groups
- Inclusion, exclusion, and missing-data rules
- Expected counts in each group after QC
- Key covariates and their coding
- Ancestry composition and planned structure adjustment
- Known or possible relatives and the analysis policy for them
- Secondary analyses that require preserved subgroups
The Genome-wide Association Analysis Service can be planned against this analysis population. The broader population genomics study-design guide discusses sequencing depth and platform choice; this page remains focused on SNP-array submission.
Size for the Final Cohort
Power calculations should use the number of analyzable samples, not the shipment total. Allow for samples that may fail DNA QC, produce low call completeness, show identity conflicts, or require exclusion because of relatedness, ancestry, or missing phenotype data. Avoid adding one arbitrary attrition percentage to every group. Historical evidence about the specific sample source and extraction method is more useful.
| Quantity | Definition | Why it belongs in the plan |
| Available samples | All records that could be considered | Shows the recruitment or collection pool |
| Eligible samples | Records meeting biological and metadata rules | Defines the intended analysis population |
| Submitted samples | Samples placed on plates, including controls and replicates | Determines laboratory scope |
| Expected QC-passing samples | Submitted biological samples likely to remain | Drives the realistic power calculation |
| Final independent sample count | QC-passing samples after relatedness policy | Drives models that require unrelated individuals |
Recalculate power under plausible group imbalance and subgroup analyses. If ancestry groups will be analyzed separately, total N can hide insufficient effective sample size in one stratum. The PCA analysis service and Population Structure Analysis Service can support cohort characterization, but these analyses do not replace balanced recruitment.
Balance Before You Randomize
Pure random assignment can still produce imbalanced plates by chance, especially for small groups. Use constrained or stratified randomization: define the factors that must be balanced, allocate them across plates, then randomize within the permitted positions. Keep the number of constraints manageable so the layout remains operational.
High-priority balance factors usually include the primary phenotype group, genetic ancestry or sampling population, study site, extraction batch, sex when relevant to the analysis, and sample quality category. Lower-priority factors can be monitored without forcing exact equality.
| Factor | Good design | Risky design |
| Primary comparison | Each plate contains a similar group ratio | One group is concentrated on one plate or processing wave |
| Ancestry or population | Major groups are represented across batches | Population and batch are nearly identical variables |
| Collection site | Sites are distributed or bridged | Each site is processed in a separate period |
| Extraction batch | Batches are mixed across plates when feasible | One extraction batch fills an entire plate |
| Input quality | Quality range is distributed | Lower-quality DNA is concentrated in one plate region |
| Controls and replicates | Positions span plates and runs | All controls occupy one plate or adjacent wells |
Figure 2. Constrained randomization preserves group balance while distributing technical factors across plates.
The QC metrics and batch-effects resource provides downstream checks, including missingness, heterozygosity, and plate-pattern review. Good analysis begins with a layout that makes those checks interpretable.
Plate Layout Needs Rules
A plate map is a controlled data object, not a convenience spreadsheet. Each occupied well should link to one unique sample ID, specimen ID where applicable, study group, site, extraction batch, DNA concentration, normalization status, and any control or replicate designation. Never encode sensitive personal information in plate IDs or sample labels.
Use these operational rules:
- Reserve and label required control positions before assigning study samples.
- Distribute study groups and major covariates across rows, columns, plates, and processing waves.
- Avoid placing all scarce or irreplaceable samples in one plate or one run.
- Separate intentional replicates from accidental duplicate IDs.
- Preserve the randomization seed, algorithm or allocation record, and any manual changes.
- Lock the final map with a version number before shipment and record all post-lock substitutions.
Edge and row effects vary by platform and process, so the design should not assume that any position is inherently safe. Distribution is the protective principle. If the laboratory specifies fixed control wells or plate constraints, incorporate them before randomization rather than moving samples after the map is finalized.
Extraction Batches Can Confound
DNA extraction is part of the technical history even when it occurred months before genotyping. Collection site, source material, storage duration, extraction kit, operator, plate, and date can influence DNA quantity and quality. If one study group was extracted with one method and another group with a different method, the downstream array signal may carry that difference.
Inventory extraction metadata before assigning plates. Where practical, mix extraction batches across array plates while maintaining traceability. If mixing is impossible because the samples arrive in separate waves, include bridge samples and avoid changing multiple process variables at the same boundary. A new extraction method, laboratory, array version, and analysis pipeline introduced together creates an attribution problem.
The DNA sample suitability guide for population genomics helps structure information on source material, storage, extraction, and low-input samples. Exact acceptance criteria should be confirmed with the selected array workflow before samples are shipped.
Controls and Replicates Have Jobs
Controls are most useful when each has a stated question. A duplicate can estimate genotype consistency, but it does not measure contamination unless the comparison is designed for that purpose. A bridge sample can connect processing waves, but only if it is stable, available in sufficient quantity, and tracked as the same biological source.
| Control type | Question answered | Placement principle |
| Technical replicate | Are genotypes reproducible across wells, plates, or runs? | Place across the factor being tested |
| Bridge sample | Are batches or project waves comparable? | Repeat across every relevant batch boundary |
| Known reference sample | Does calling agree with an established genotype? | Include according to platform and project plan |
| Negative control | Is there evidence of contamination or non-specific signal? | Use when supported by the laboratory workflow |
| Blind duplicate | Can sample handling and identity be checked end to end? | Conceal relationship from routine processing where appropriate |
Technical replicates are especially valuable during a new array configuration, a multicenter handoff, or a long project with multiple processing waves. They consume capacity, so choose them to test the highest-risk boundaries rather than repeating samples without a rationale. Laurie and colleagues emphasized balancing and randomizing study factors, while later work has shown that batch composition can affect genotype calling and downstream results.
Multi-Site Cohorts Need Bridges
Consortia and longitudinal cohorts often receive samples by site or timepoint. The design should prevent site, time, storage, extraction, and genotyping batch from becoming one composite variable. A phased project needs an explicit bridge strategy and a single metadata model across waves.
Practical safeguards include:
- A shared sample-manifest specification used by every site
- Central validation of identifiers and allowed values before plate assignment
- Bridge samples or replicates across processing waves
- Stable array version and genotype-calling workflow where feasible
- Documented change control when reagents, instruments, manifests, or software change
- Batch-aware QC summaries before cohorts are merged
The Human 85K SNP Genotyping Array Service can incorporate cohort review, sample normalization, plate strategy, genotype calling, and analysis-ready deliverables. For projects using another configuration or species, start with the broader SNP Genotyping Service.
Metadata Must Match Samples
Genotyping quality is not useful if sample identity and phenotype metadata cannot be reconciled. Use one authoritative sample key and keep separate fields for participant, specimen, extraction, plate, and analysis IDs. Do not overwrite identifiers during normalization or replace missing values with ambiguous codes.
Minimum manifest fields include:
- Unique analysis sample ID and source specimen ID
- Plate ID, well, and processing wave
- Study group, population or ancestry metadata, and collection site
- Sample source, extraction method, extraction batch, storage, and buffer
- DNA concentration, volume, normalization status, and QC flags
- Replicate or control relationship
- Genome build and requested output formats
- Primary phenotype availability and covariate completeness status
Run automated checks for duplicate IDs, missing wells, impossible category values, sample-to-phenotype mismatches, and inconsistent replicate labels. Freeze a submission version, then create a separate change log rather than editing the original silently.
Figure 3. A versioned sample key connects every plate position and QC decision to the final genotype dataset.
Pre-Submission Checklist
The project is ready for technical review when the following items are complete:
- Primary and secondary analyses are named.
- Cohort counts are reconciled from available to expected QC-passing samples.
- Group balance, ancestry composition, relatedness, and covariates are documented.
- Sample source, extraction batch, storage, concentration, volume, and buffer are recorded.
- Plate constraints, reserved wells, controls, replicates, and bridge samples are defined.
- A constrained randomization method and change-control record are available.
- The exact array manifest and genome build are confirmed.
- Sample-level and marker-level QC outputs are specified.
- Imputation, PCA, relatedness, ROH, or GWAS deliverables are included when needed.
- The submission manifest has passed identifier and allowed-value checks.
Do not ship samples until the laboratory has confirmed the specification, logistics, and final plate plan. CD Genomics can review the array workflow, downstream bioinformatics scope, and handoff requirements before submission.
Stress-Test the Layout Before Submission
Use a hypothetical failure scenario to audit the design. Suppose cases were recruited at one site during the first year and controls at another site during the second year. If samples are processed in arrival order, phenotype, site, storage duration, extraction batch, and plate may become nearly interchangeable variables. A later association between genotype quality and case status would then be difficult to interpret. Software covariates cannot reliably recover information that the physical design never separated.
Build a cross-tabulation of every major biological factor against plate, row, column, extraction batch, site, and processing wave. Flag empty cells, very small strata, and perfect or near-perfect alignment. Then create a constrained randomization that distributes the primary groups while respecting operational requirements such as reserved control wells and sample-volume limits. Keep the random seed, input manifest, exclusions, and final plate map. If balancing is impossible, document the limitation before processing and decide whether bridging samples, staged analysis, or additional recruitment can reduce the ambiguity.
After data return, troubleshoot patterns at the correct level. Plate-wide shifts point toward processing or reagent effects; row or column patterns suggest liquid-handling or position effects; site-specific missingness may reflect collection or extraction; sex-chromosome discordance, unexpected duplicates, or relatedness conflicts can indicate identity or metadata problems. Review these signals alongside ancestry and study group before excluding samples. A QC threshold from a published cohort is a practical starting point, not a universal requirement, because platforms, populations, sample sources, and analytical goals differ.
The design is ready when every sample can be traced from source specimen to analysis identifier, every planned comparison is represented across technical units where feasible, and every unavoidable imbalance is named. Genotyping can support association and population analyses, but it cannot make an observational association causal. Important findings may need independent replication, orthogonal genotyping, or functional validation. Technical replicates are most valuable when they cross a suspected risk boundary; adding many replicates within one uncomplicated plate may add less information than a smaller set placed across sites, batches, or waves.
Frequently Asked Questions
Place both groups on every plate in a similar ratio wherever practical, while also balancing ancestry, site, extraction batch, sex, and input quality. Avoid layouts in which phenotype group and plate are the same variable.
Yes, when feasible and operationally safe, because mixing helps separate extraction effects from array-plate effects. Keep full traceability and use bridge samples when batches cannot be mixed.
Replicates are most useful across high-risk boundaries such as plates, runs, sites, array versions, or processing waves. Define the expected comparison and acceptance rule before selecting the samples.
Analysis can identify and sometimes reduce batch-related variation, but it cannot reliably solve perfect confounding. Balanced design and traceable metadata remain the primary controls.
Maintain a common manifest, stable workflow, and bridge-sample strategy. Review each wave independently, compare shared controls and QC distributions, and document every process change before merging data.
References
- Truong VQ, Woerner JA, Cherlin TA, et al. Quality Control Procedures for Genome-Wide Association Studies. Current Protocols. 2022;2(11):e603. doi:10.1002/cpz1.603.
- Yang H-C, Kwok P-Y, Li L-H, et al. The Taiwan Precision Medicine Initiative provides a cohort for large-scale studies. Nature. 2025;648(8092):117-127. doi:10.1038/s41586-025-09680-x.
- Sharma S, Nagar SD, Pemu P, et al. Genetic ancestry and population structure in the All of Us Research Program cohort. Nature Communications. 2025;16(1):4123. doi:10.1038/s41467-025-59351-8.
- Uffelmann E, Huang QQ, Munung NS, et al. Genome-wide association studies. Nature Reviews Methods Primers. 2021;1(1):59. doi:10.1038/s43586-021-00056-9.
- Seo S, Park K, Lee JJ, et al. SNP genotype calling and quality control for multi-batch-based studies. Genes & Genomics. 2019;41(8):927-939. doi:10.1007/s13258-019-00827-5.
- Johnson R, Ding Y, Bhattacharya A, et al. The UCLA ATLAS Community Health Initiative: Promoting precision health research in a diverse biobank. Cell Genomics. 2023;3(1):100243. doi:10.1016/j.xgen.2022.100243.
- Laurie CC, Doheny KF, Mirel DB, et al. Quality control and quality assurance in genotypic data for genome-wide association studies. Genetic Epidemiology. 2010;34(6):591-602. doi:10.1002/gepi.20516.
For research purposes only. The information and services described here are not intended for clinical diagnosis, therapeutic decisions, or personal health assessment.