Generate standardized human genotype data across large cohorts with an approximately 85K-marker SNP array configuration. CD Genomics supports population genomics, GWAS-oriented studies, cohort stratification, and downstream statistical genetics with integrated sample QC and analysis-ready deliverables.
Large human cohorts require more than a collection of individual genotype files. They need a consistent marker framework, traceable sample handling, cohort-aware quality control, and data structures that can move directly into statistical genetics. CD Genomics provides a human SNP genotyping array service built around an approximately 85K-marker configuration for research teams that need standardized genome-wide genotypes across tens, hundreds, or thousands of samples.
The service connects study-design review, DNA quality assessment, array processing, genotype calling, data harmonization, and optional downstream analysis. Rather than treating genotyping as an isolated laboratory step, we organize the workflow around the research question: identifying population structure, detecting unexpected relatedness, studying runs of homozygosity, preparing a GWAS dataset, or generating array data suitable for genotype imputation.
Researchers who need a different species, another marker density, or a targeted method can review our broader SNP Genotyping Service.
Figure 1: Human 85K SNP array service overview from cohort DNA to standardized genotypes and population-genomics outputs.
An approximately 85K-marker array occupies a practical middle ground between small targeted panels and higher-density or sequencing-based approaches. Because the loci are predefined, the same marker framework can be applied across the cohort, supporting consistent genotype representation and efficient data management. This is particularly useful when the study priority is comparison across many individuals rather than open-ended discovery of every genomic variant.
The array is best suited to research projects that need:
The array is not designed for comprehensive novel-variant discovery, direct detection of all rare variants, or complete characterization of structural variation. Projects centered on these questions may be better served by sequencing.
The service uses an approximately 85K-marker human SNP array configuration with genome-wide coverage that can include chromosomes 1–22 and X. The final marker count and composition are confirmed during project setup because available content can vary with the selected array configuration, reference build, and any technically feasible marker supplementation.
Marker density alone does not determine whether an array is appropriate. Study population, linkage disequilibrium patterns, ancestry composition, phenotype definition, relatedness, target allele frequency, and the intended downstream analysis all influence study utility. Our project review therefore begins with the research objective rather than a one-size-fits-all marker claim.
| Study-Design Input | Why It Matters |
| Cohort size and sampling scheme | Affects statistical power, plate layout, batch planning, and the handling of related participants. |
| Population and ancestry composition | Influences marker informativeness, population-structure assessment, and reference-panel selection for imputation. |
| Research endpoint | Determines whether the workflow should prioritize PCA, relatedness, ROH, association testing, or imputation preparation. |
| Reference assembly | Controls chromosome coordinates, allele representation, annotation, and downstream compatibility. |
| Phenotype and covariate data | Supports association-ready sample mapping and helps identify metadata inconsistencies before analysis. |
| Requested marker additions | Requires probe and platform feasibility review; supplementation is not guaranteed for every locus. |
Project-specific marker supplementation or an alternative density can be evaluated where the platform and probe design allow. This review does not imply unrestricted custom array design.
Figure 3: Illustrative marker-coverage concept for an approximately 85K human SNP array; final content depends on project configuration.
We define the cohort, study endpoint, sample layout, genome build, marker configuration, metadata structure, optional analyses, and final deliverables.
Sample identifiers and accompanying metadata are checked for completeness, uniqueness, and compatibility with the proposed plate map.
DNA quantity, purity, integrity, and concentration consistency are evaluated using methods appropriate to the submitted material and project scope. Samples are normalized for array processing when applicable.
Samples undergo the platform-specific amplification, hybridization, washing, staining, scanning, and signal collection workflow associated with the finalized array configuration.
Genotype clusters are evaluated, sample- and marker-level metrics are calculated, and potential plate or batch effects are reviewed before data release.
Approved samples and markers are harmonized into analysis-ready files. Optional modules may include population structure, relatedness, ROH, GWAS preparation, or imputation preparation.
Figure 2: End-to-end workflow for cohort-scale human SNP array genotyping.
Purified human genomic DNA is the standard input. Exact acceptance criteria are confirmed after review of the array configuration, cohort size, extraction method, and available sample volume. This prevents fixed specifications from being applied to projects with different matrices or operational constraints.
| Input | Preferred Project Information | Review Points |
| Purified genomic DNA | Unique sample ID, concentration, volume, extraction method, source material, and storage buffer | Quantity, purity, integrity, concentration consistency, and compatibility with array processing |
| Biospecimens requiring extraction | Sample type, collection method, storage conditions, shipping history, and available volume | Extraction feasibility must be confirmed before acceptance; purified DNA may be recommended |
| Sample manifest | One row per sample with unique IDs and non-identifying study metadata | Duplicate IDs, missing fields, plate-map compatibility, and data-transfer requirements |
| Phenotype and covariate table | Variables required for the planned research analysis | Coding consistency, missingness, allowed values, and sample-ID correspondence |
| Requested marker list | rsID and/or genomic coordinate, reference build, alleles, and priority | Platform compatibility, probe feasibility, redundancy, and final configuration impact |
Please do not ship samples until the project team has confirmed the sample specification and logistics.
Array processing produces genotype calls, but a population-genomics dataset also requires cohort-level review. The QC plan is agreed before analysis so that filters are appropriate for the population, study design, array content, and intended endpoint. Fixed thresholds are not imposed without considering these factors.
Researchers who need dedicated downstream interpretation can combine the service with Population Structure Analysis or Genome-wide Association Analysis.
Deliverables are finalized during project design and may include:
.pgen, .pvar, .psam) or PLINK 1 binary formats (.bed, .bim, .fam)The final delivery package is designed around traceability: researchers can see which samples and markers were retained, flagged, or excluded and which decisions were applied before downstream analysis.
Use genome-wide genotype patterns to characterize cohort heterogeneity, identify genetic clusters, visualize ancestry gradients, and define covariates for downstream analyses.
Detect duplicated samples, unexpected close relatives, and cryptic relatedness that may affect association tests or cohort composition.
Measure long homozygous segments and compare ROH burden or distribution across research groups, subject to marker density and QC design.
Prepare standardized genotypes, QC metrics, population covariates, and sample mappings for association analysis of complex traits or disease-associated variation in research cohorts.
Harmonize build, chromosome, coordinate, strand, allele, and marker identifiers before phasing and imputation. Imputation performance depends on marker overlap, data quality, ancestry match, and reference-panel composition.
Use PCA, relatedness, heterozygosity, and missingness information to select balanced or independent subsets for follow-up experiments.
The demo below uses simulated data and is intended to explain potential deliverables rather than report service performance. It combines a three-cluster genotype intensity plot, a sample-level completeness distribution, a PCA scatter plot, and a relatedness heatmap. Any review boundary shown is illustrative and is not a CD Genomics guarantee or universal acceptance threshold.
Figure 4: Illustrative simulated genotype-calling, QC, population-structure, and relatedness outputs; values are not service specifications.
The Taiwan Precision Medicine Initiative provides a cohort for large-scale studies
Journal: Nature
Published: 2025
Yang H-C, Kwok P-Y, Li L-H, et al. The Taiwan Precision Medicine Initiative provides a cohort for large-scale studies. Nature. 2025;648(8092):117–127.
Large population cohorts can be limited when the available SNP content and reference resources do not adequately represent the study population. The Taiwan Precision Medicine Initiative established a large Han Chinese research cohort and used population-optimized SNP arrays to support standardized genotyping and downstream genetic analyses.
The study used two customized SNP arrays, TPMv1 and TPMv2. Genotyping was organized in batches, and the investigators applied sample- and marker-level QC, genotype harmonization, relatedness analysis, population-structure analysis, GWAS workflows, and phasing/imputation validation. For the imputation assessment, 6,000 genotyped variants on chromosomes 5, 13, and 18 were masked in 1,000 individuals and compared with their imputed genotypes.
After the reported QC measures, the study had genotype data from 165,596 individuals processed with TPMv1 and 321,360 with TPMv2, totaling 486,956 participants with matching research records. The masked-variant validation reported an average correlation of 0.906 between imputed and observed genotypes, and 96.3% of masked SNPs had an INFO score above 0.7. The investigators also used shared, QC-passed SNPs for PCA, ancestry characterization, relatedness assessment, and association analyses.
Figure 5: Original summary of the published Taiwan Precision Medicine Initiative SNP-array workflow and reported cohort-level analyses. Source: Yang et al. (2025).
The study shows how standardized SNP-array genotyping can be scaled across a very large cohort when marker selection, batch normalization, sample QC, marker QC, population structure, relatedness, and imputation are treated as connected parts of one research workflow. For service planning, the transferable lesson is the importance of cohort-aware design and QC—not the adoption of a universal marker count or fixed performance threshold.
The most appropriate method depends on cohort size, ancestry, allele-frequency range, discovery requirements, reference resources, analysis objectives, and budget. Method selection should begin with the research question rather than a preferred technology.
| Method | Best Fit | Main Strength | Important Limitation |
| Approximately 85K human SNP array | Large human cohorts focused on predefined genome-wide markers | Consistent loci, efficient cohort processing, and direct compatibility with common genotype analyses | Does not comprehensively discover novel, rare, or structural variants |
| Genotyping By Sequencing | Projects needing reduced-representation sequencing or flexible marker discovery, especially outside standardized human-array settings | Combines variant discovery and genotyping without a fixed array | Locus representation and missingness can vary with library and sequencing factors |
| Low-pass whole-genome sequencing | Population-scale projects seeking broader genome representation with imputation | Samples the genome beyond a predefined array and can support imputed genotypes | Requires sequencing, imputation, and additional computational QC; utility depends on coverage and reference panel |
| Whole Genome Re-sequencing | Projects prioritizing broad variant discovery and richer genomic characterization | Most comprehensive option among these methods for sequence-based discovery | Higher data volume, analytical complexity, and project cost than a fixed SNP array |
Confidence in a population-scale array project comes from decisions that remain traceable from marker configuration and sample layout through cohort-level delivery. CD Genomics connects laboratory processing, quality control, and downstream requirements so the genotype dataset is designed for the research question rather than treated as an isolated assay output.
Cohort-aware study planning: marker configuration, sample layout, metadata, QC criteria, and statistical endpoints are connected before processing begins. This reduces the risk of producing technically valid genotypes that do not fit the planned analysis.
Any sample, marker, output, timing, or performance specification not confirmed for the individual project remains subject to feasibility review. This transparency helps identify method limitations before they become unexplained gaps in a large cohort.
No. "85K" describes an approximate configuration. The final count can vary with the selected array version, genome build, QC status, and any technically feasible marker supplementation. The confirmed marker manifest is part of project setup.
Project-specific supplementation can be evaluated when probe design and the selected platform support the requested loci. Please provide rsIDs or genomic coordinates, reference build, alleles, and marker priorities. Feasibility is reviewed before the configuration is finalized.
It can support GWAS-oriented research, but suitability depends on cohort size, effect size, allele frequency, phenotype quality, ancestry composition, marker coverage, imputation strategy, and statistical design. A power and design review is recommended before genotyping.
Yes, the workflow can prepare data for imputation by checking reference build, chromosome coordinates, strand, alleles, duplicates, missingness, and marker overlap. Imputation accuracy is not determined by the array alone; ancestry match and reference-panel composition are also important.
Thresholds are set according to the study population, analysis endpoint, array configuration, and cohort distributions. The project plan can define initial review criteria, while final decisions are documented in the QC report.
Purified genomic DNA is the standard input. Required quantity, concentration, purity, integrity, volume, and buffer are confirmed after the array and cohort scope are reviewed. Samples should not be shipped before specifications are approved.
Batch planning begins with plate layout and sample randomization or balancing where appropriate. During analysis, sample and marker metrics are reviewed by plate and batch, and genotype consistency, missingness, intensity patterns, and population covariates can be evaluated for systematic differences.
No. An SNP array assays predefined loci. If novel, rare, or structural variant discovery is central to the research question, low-pass or standard-depth whole-genome sequencing may be more appropriate.
Typical deliveries include genotype files, sample and marker QC tables, a manifest, a data dictionary, and a methods summary. PLINK, VCF, PCA, relatedness, ROH, GWAS-ready, or imputation-preparation outputs are included when specified in the project scope.
The schedule depends on sample count, input quality, array availability, marker configuration, repeat requirements, downstream analysis, and delivery format. A project-specific timeline is provided after technical review.
Use the SNP Genotyping Service when the project requires a different species, marker density, or genotyping technology. Consider Genotyping By Sequencing or Whole Genome Re-sequencing when sequence-based marker discovery is required. For downstream cohort analysis, explore Population Structure Analysis and Genome-wide Association Analysis.
Next step: prepare the cohort size, sample source, ancestry composition, reference build, phenotype structure, requested markers, intended analyses, and preferred delivery formats for a project-design review.
References
For research purposes only, not intended for clinical diagnosis, treatment, or individual health assessments.