Global Diversity Array and Infinium Arrays: High-Throughput Genotyping for Human and Agricultural Genomics

Fixed-Content Arrays — The Workhorse of Large-Scale Genotyping

Fixed-content SNP arrays occupy a distinct niche in the genotyping technology landscape that neither whole-genome sequencing nor reduced-representation sequencing has displaced — and understanding why reveals the logic that should drive platform selection decisions for core facilities and breeding programs alike.

The Illumina Infinium BeadChip platform is the dominant technology in this space. Each BeadChip is a silica wafer etched with microwells containing oligonucleotide-derivatized beads, each bead coated with tens of thousands of copies of a locus-specific 50-mer probe. Two assay chemistries coexist: Infinium I uses two bead types per SNP (one per allele) and discriminates alleles by single-base extension with differentially labeled dideoxynucleotides; Infinium II uses a single bead type per SNP with two-color staining after extension, trading slightly lower call rates for higher multiplexing density. A single BeadChip processes 8, 24, or 48 samples depending on the array format, with the number of markers per sample ranging from approximately 50,000 on low-density agricultural arrays to nearly 2 million on the Global Diversity Array with Enhanced PGx content.

The mid-2025 release of Infinium EX chemistry reduced DNA input requirements and shortened processing time compared to the legacy Infinium HD chemistry. EX-powered products include the Global Screening Array v4 (GSA v4, 48 samples per chip, ~650,000 markers with up to 50,000 custom add-on slots) and the Global Clinical Research Array (GCRA, 24 samples per chip, ~1.2 million markers curated from leading research databases including ClinVar, PharmGKB, and ClinGen). Both have enhanced Pharmacogenetics (PGx) variants with targeted gene amplification for challenging loci such as CYP2D6 and CYP2B6 where pseudogene interference complicates standard genotyping.

The fundamental advantage of fixed-content arrays — and the reason they persist despite the declining cost of sequencing — is data consistency. Every sample run on a given array design is interrogated at exactly the same set of markers, producing genotype calls that are directly comparable across batches, laboratories, and years without imputation alignment or batch-effect correction. For a core facility that processes samples from a longitudinal cohort study across multiple years, or a breeding program that accumulates genotype data across generations of selection candidates, this comparability is not a convenience — it is an operational requirement. A genotype at rs429358 on a GDA run in 2025 is directly comparable to a genotype at the same SNP on a GDA run in 2028, with no imputation step, no reference panel dependency, and no batch-effect confounders to model. For a broader overview of the genotyping technology landscape including sequencing-based approaches such as GBS and ddRAD-seq, see CD Genomics' Genotyping and Genetic Diversity Services overview.

Figure 1: Infinium BeadChip Technology — From DNA to Genotype Calls via Single-Base ExtensionFigure 1: Infinium BeadChip Technology — From DNA to Genotype Calls via Single-Base Extension

Human Genotyping Arrays — GDA and Its Derivatives

The Infinium Global Diversity Array (GDA) is the current flagship human genotyping array, carrying approximately 1.8 million markers on the base configuration and expandable to approximately 2 million with booster content modules. The SNP backbone was designed using whole-genome sequence data from the Consortium on Asthma among African-ancestry Populations in the Americas (CAAPA) and the Population Architecture using Genomics and Epidemiology (PAGE) study — over 1,000 whole genomes of African ancestry informed marker selection, addressing the persistent Eurocentric bias that limited the performance of earlier arrays in non-European populations. The GDA was selected as the genotyping platform for the NIH All of Us precision medicine program, which alone represents one of the largest single-platform genotyping commitments in history.

The GDA ecosystem has expanded through booster content add-ons — secondary bead pools manufactured on the same BeadChip that supplement the core marker set with domain-specific content. The Enhanced PGx booster adds approximately 108,000 markers covering over 44,000 ADME-related variants and 2,000 pharmacogenetic targets across more than 1,700 star alleles, with targeted gene amplification for CYP2D6 to disambiguate pseudogene interference. The Polygenic Risk Score Content booster (v1.0) adds markers selected for pan-ethnic PRS determination, bringing total marker count to approximately 2 million. The Carrier Screening Content booster (v2.0) adds approximately 45,000 markers across more than 600 clinically significant genes. The Cytogenetics booster optimizes coverage for CNV detection with probes targeting over 4,800 clinically relevant genes. A core facility offering the GDA can configure the booster stack to match the research programs it serves — PGx for a pharmacogenomics cohort, PRS for a population-health study, Cytogenetics for a developmental-disorder registry — without changing the base platform, consumables supply chain, or analysis pipeline.

The GSA v4, powered by Infinium EX chemistry, serves a different niche: higher throughput (48 samples per BeadChip) at lower per-marker cost (~650,000 markers) with custom add-on capacity, making it the economical choice for biobank-scale projects where per-sample cost dominates platform selection. The Asian Screening Array (ASA) and the NeuroBooster Array (NBA) — the latter built on the GDA backbone with approximately 95,000 neurological-disease-specific markers — illustrate the platform's adaptability: the same BeadChip manufacturing process, the same iScan reader, the same GenomeStudio analysis software, but radically different biological content tuned to specific populations and disease areas.

When should a human genetics study choose an array over sequencing? The answer turns on three variables: whether the markers of interest are known in advance, whether samples will be collected and analyzed over an extended period, and whether the budget favors per-sample cost at scale or per-marker informational content. A study that needs to genotype 10,000 samples at a defined set of 650,000 SNPs for PRS calculation, with samples arriving over three years, is an array project. A study that needs to discover novel rare variants in 500 samples with a specific phenotype is a sequencing project. The cost crossover — the sample number above which arrays become cheaper per unit of useful genetic information than sequencing — depends on sequencing depth, array density, and the imputation reference panel available, but as a practical benchmark: at current pricing, an array-based GWAS on 5,000 samples costs less than 0.5× lcWGS on the same cohort, and produces genotype calls at known, well-validated markers without the imputation uncertainty that accompanies low-coverage sequencing genotypes.

A concrete example illustrates the booster-configuration logic in practice. A research hospital running a pharmacogenomics cohort of 3,000 patients on antidepressant therapy needs both genome-wide SNP data for population stratification and targeted PGx variants for CYP2D6 and CYP2C19 metabolizer status prediction. Running the base GDA followed by a separate targeted genotyping panel for PGx variants adds per-sample cost and a data-merging step. Running the GDA with Enhanced PGx booster — a single BeadChip workflow delivering 1.93 million markers including all required PGx content — eliminates the merging step and reduces total per-sample cost by approximately 25 to 35 percent compared to the two-platform approach. The booster decision is not about whether the content exists; it is about whether the incremental per-sample cost of the booster is less than the cost of running a separate assay for the same markers — a calculation that depends on sample number, the size of the booster content relative to a standalone targeted panel, and the personnel cost of merging data from two platforms. SNP microarray services supporting the full GDA family with booster configuration and GenomeStudio-based genotype calling are available.

Figure 2: Human vs. Agricultural Genotyping Arrays — Marker Density, Sample Format, and Application DomainsFigure 2: Human vs. Agricultural Genotyping Arrays — Marker Density, Sample Format, and Application Domains

Agricultural Arrays — Species-Specific Genotyping at Scale

Agricultural genotyping arrays solve a fundamentally different problem than human arrays. Human arrays are designed to capture population-wide common variation for complex-trait GWAS and PRS — markers are selected to tag haplotype blocks across diverse ancestries. Agricultural arrays are designed for genomic prediction: estimating breeding values from genome-wide marker genotypes, where the goal is not identifying causal variants but accurately capturing genome-wide relationships among selection candidates. This difference in purpose drives every design choice.

Species-specific commercial arrays include the BovineSNP50 (approximately 50,000 markers), BovineHD (approximately 777,000 markers), PorcineSNP60 (approximately 65,000 markers), MaizeSNP50 (approximately 50,000 markers), and various wheat arrays ranging from 9,000 to 80,000 markers depending on consortium design. Marker density on agricultural arrays is consistently lower than on human arrays for the same cost tier because agricultural species have not benefited from the massive reference-panel investments (1000 Genomes, TOPMed, gnomAD) that enabled human array optimization. Agricultural arrays are designed by consortia — the USDA, Wageningen University, the International Maize and Wheat Improvement Center (CIMMYT), and various national breeding programs — and the markers reflect the priorities of those consortia: uniform genome coverage for genomic prediction, inclusion of known functional variants (disease resistance genes, major QTLs, parentage SNPs), and compatibility across breeds, lines, and populations within a species.

A notable 2025 development is the IMAGE001 multispecies array, a 10,000-SNP panel covering six livestock species (cattle, sheep, goat, horse, pig, and chicken) on a single BeadChip format. Designed specifically for gene bank collections, IMAGE001 includes functional variants, mitochondrial DNA markers, sex-chromosome markers, and disease- and trait-related SNPs. For a gene bank managing cryopreserved material across multiple livestock species, IMAGE001 replaces six separate species-specific genotyping workflows with one — a practical advance that reduces cost, simplifies laboratory logistics, and standardizes QC across species. This is the kind of innovation that matters to a core facility but is invisible to most journal readers.

Genomic selection — the primary application of agricultural SNP arrays — uses genome-wide markers to predict breeding values without phenotyping every individual. GBLUP (genomic best linear unbiased prediction) remains the industry standard for its balance of accuracy and computational efficiency, while Bayesian methods (BayesA, BayesB, BayesR) and emerging deep-learning frameworks (Gsformer, a CNN-self-attention architecture published in 2026) offer incrementally higher accuracy at significantly higher computational cost. The key parameter for array users is not the specific prediction algorithm but the marker density: prediction accuracy saturates at a species- and trait-dependent marker count, and adding markers beyond that saturation point costs money without improving selection decisions. For dairy cattle, accuracy plateaus at approximately 50,000 to 80,000 markers; for highly diverse crops with rapid LD decay such as maize, 10,000 to 30,000 markers may be adequate. A core facility advising a breeding program should recommend the lowest-density array that achieves the program's accuracy targets, not the highest-density array the budget can accommodate.

Marker-assisted backcrossing — using a small set of diagnostic markers to track known QTLs or transgenes across backcross generations — represents the opposite end of the agricultural genotyping spectrum: few markers, many samples, rapid turnaround. For 1 to 10 markers across thousands of samples, targeted assays (KASP or TaqMan, discussed in our article on Microsatellite and SNP Genotyping Services) are more cost-effective than running a 50,000-SNP array and discarding 49,990 genotypes per sample.

Figure 3: Genotyping Technology Decision Matrix — Arrays vs. GBS vs. lcWGS across Sample Size and Marker RequirementsFigure 3: Genotyping Technology Decision Matrix — Arrays vs. GBS vs. lcWGS across Sample Size and Marker Requirements

Custom Array Design — When Off-the-Shelf Is Not Enough

The fixed-content SNP arrays described above are products of consortia, designed for well-resourced species with large user communities. A plant geneticist working on European white oak, an ecologist studying a non-model fish species, or a breeder developing a novel crop variety may find that no commercial array exists for their species — or that existing arrays were designed for populations genetically distant from their study material and perform poorly due to SNP ascertainment bias. Custom array design fills this gap, but the path from a set of candidate SNPs to a validated, production-ready BeadChip is more demanding than the commercial array workflow.

The custom array design cycle proceeds in four stages. Stage 1 — SNP discovery — typically uses whole-genome resequencing or reduced-representation sequencing (GBS, ddRAD-seq) of 50 to 200 individuals representing the target populations to identify polymorphic loci. Stage 2 — SNP selection and in silico scoring — filters candidate SNPs through Illumina's proprietary Assay Design Tool (ADT), which scores each SNP for predicted Infinium assay performance based on flanking sequence uniqueness, GC content, and the presence of secondary variants in the probe-binding region. SNPs receiving an ADT score below 0.6 are typically excluded, as they are unlikely to produce reliable genotype clusters regardless of how well-designed the rest of the assay may be. The ADT scoring process is not a single-pass filter. A common core-facility workflow iterates: submit the full candidate SNP list, receive ADT scores, remove SNPs scoring below 0.4 outright, then for SNPs scoring 0.4 to 0.6 examine the flanking sequence for secondary variants within 60 bp of the target SNP that can be avoided by shifting the probe design window, re-submit the adjusted probe sequences for re-scoring, and retain SNPs that improve to 0.6 or above on the second pass. This iterative rescue step recovers 10 to 15 percent of SNPs that would be discarded by a single-pass ADT threshold, which is material when the starting pool of candidates is limited — a common situation in non-model species where initial SNP discovery yields 20,000 to 40,000 variants rather than the 100,000-plus typical of model organisms. A typical custom array design starts with 50,000 to 100,000 candidate SNPs, of which 60 to 80 percent pass ADT scoring, and 30,000 to 60,000 are selected for the final array based on minor allele frequency, genomic distribution, and functional annotation.

Stage 3 — manufacturing — is performed by Illumina; turnaround time from design submission to BeadChip shipment is typically 8 to 12 weeks. Stage 4 — validation — is where custom arrays succeed or fail. A foundational study that designed a custom Infinium array for European white oaks (Quercus petraea and Q. robur) illustrates the real-world challenges: from 7,913 candidate SNPs submitted to design, genotyping success rate was 80.4 percent in mapping populations but dropped to 54.8 percent in natural populations, with approximately 25 percent of successfully genotyped SNPs showing cluster compression indicative of paralogue interference. The authors estimated that 15 to 20 percent of SNPs on the final array were unreliable for population-genetic inference, and recommended genotyping a diverse pilot set of 48 to 96 samples before committing the full study cohort to a custom array — a recommendation that core facilities should treat as a minimum standard, not an optional step.

The economic decision for custom arrays turns on sample throughput. A custom 50,000-SNP BeadChip with a minimum order of 288 samples (one 24-sample chip format, 12 chips) has a higher per-sample cost than the equivalent commercial array at scale, but for a species with no commercial array — and no prospect of one — the choice is not between custom and commercial but between a custom array and an alternative genotyping technology. If the alternative is GBS at 10,000 to 20,000 SNPs per sample with 20 to 40 percent missing data requiring imputation, and the study requires marker consistency across 2,000 samples collected over five years, the custom array may be the more economical choice despite higher upfront design cost because it eliminates per-batch imputation, batch-effect correction, and the data-management overhead of merging VCF files from multiple GBS runs.

Figure 4: Custom SNP Array Design Cycle — From SNP Discovery through In Silico Scoring to Pilot ValidationFigure 4: Custom SNP Array Design Cycle — From SNP Discovery through In Silico Scoring to Pilot Validation

Data Generation and QC — The Core Facility Perspective

For the scientist running the array, the laboratory workflow is mature and well-documented, but the difference between adequate and excellent genotyping data lies in the QC decisions made at each step — decisions that require judgment, not just protocol compliance.

Sample intake and DNA quality control. The standard Infinium protocol requires 200 ng of double-stranded genomic DNA at a minimum concentration of 50 ng/µL, with a 260/280 absorbance ratio between 1.8 and 2.0. Concentrations measured by fluorometry (Qubit or QuantiFluor) rather than spectrophotometry (NanoDrop) are essential because spectrophotometric readings are inflated by degraded single-stranded DNA, RNA contamination, and residual phenol — all of which will produce an acceptable NanoDrop reading and a failed array. DNA integrity — assessed by agarose gel electrophoresis or a fragment analyzer — should show a predominant band above 10 kb. Partially degraded DNA (smearing below 10 kb with intact high-molecular-weight DNA still visible) often produces acceptable call rates but elevated GenCall score variance across SNPs, particularly at loci with high GC content where probe hybridization kinetics are sensitive to fragment length.

The laboratory workflow — whole-genome amplification (overnight, isothermal), enzymatic fragmentation, precipitation and resuspension, hybridization to the BeadChip, single-base extension with labeled dideoxynucleotides, and fluorescent staining — is largely automated on liquid-handling platforms (Tecan, Beckman) with minimal hands-on time. The BeadChip is scanned on an iScan system, which uses a confocal laser to excite the Cy3 and Cy5 fluorophores incorporated during extension and records fluorescence intensity for each bead type at each marker in two color channels, producing .idat files (one per channel per sample).

Genotype calling and QC metrics. GenomeStudio 2.0 converts raw .idat intensity data into genotype calls using a proprietary clustering algorithm. For each SNP, the software plots normalized intensity in the two color channels for all samples on the BeadChip and partitions the data into three genotype clusters (AA, AB, BB). The primary QC metrics produced are:

Call Rate

Call rate — the proportion of samples with a genotype call at a given SNP, or the proportion of SNPs called for a given sample. Industry standard for a passing SNP is call rate ≥ 98 percent across samples; for a passing sample, call rate ≥ 95 percent across SNPs. A sample with call rate below 95 percent usually reflects degraded DNA, insufficient input, or a processing error at the amplification or fragmentation step, and should be excluded and re-run rather than statistically rescued.

GenCall Score

GenCall score — a per-genotype quality metric (0 to 1) reflecting the confidence of the cluster assignment. Scores above 0.7 are generally acceptable; scores between 0.5 and 0.7 flag genotypes that should be manually reviewed in the cluster plot; scores below 0.5 should be treated as missing data. GenomeStudio reports a GenCall score for every genotype call at every SNP, and the distribution of these scores across a sample is a more sensitive indicator of DNA quality than call rate alone — a sample with 99 percent call rate but 5 percent of genotypes with GenCall below 0.5 has a data-quality problem that aggregate call rate conceals.

Cluster Separation

Cluster separation — a per-SNP metric (0 to 1) reflecting how cleanly the three genotype clusters are resolved. Values above 0.8 indicate well-separated clusters; values between 0.5 and 0.8 indicate clusters that are distinguishable but overlapping; values below 0.5 indicate SNPs where genotypes cannot be reliably assigned. SNPs with cluster separation below 0.6 should be excluded or manually re-clustered with custom boundaries. Poor cluster separation is the most common marker-level QC failure on Infinium arrays and has three primary causes: (1) a secondary SNP in the probe-binding region that reduces hybridization efficiency for one allele, producing a heterozygote cluster shifted toward one homozygote; (2) paralogous sequences elsewhere in the genome that co-hybridize to the probe, adding signal intensity to both color channels and compressing all clusters toward the origin; and (3) copy number variation at the target locus that creates additional clusters (AAAA, AAAB, AABB, etc.) not modeled by the standard three-cluster algorithm.

Reproducibility

Reproducibility — assessed by including 5 to 10 percent blind duplicate samples (same DNA, independently processed from amplification through scanning) — should exceed 99.5 percent concordance. Discordant duplicate calls are enriched for SNPs with low GenCall scores and marginal cluster separation, which is why QC pipelines should process these metrics jointly rather than applying independent thresholds: a SNP with call rate 98.5 percent, cluster separation 0.65, and duplicate concordance 97 percent passes each individual filter but carries a systematic error rate that merits exclusion or manual curation.

Figure 5: Infinium Array QC Dashboard — Call Rate, GenCall Score, Cluster Separation, and Reproducibility MetricsFigure 5: Infinium Array QC Dashboard — Call Rate, GenCall Score, Cluster Separation, and Reproducibility Metrics

Cost, Throughput, and the Array-vs.-Sequencing Calculus

The cost of array-based genotyping follows an economy-of-scale curve that differs from sequencing in one critical respect: the incremental per-sample cost declines with sample number primarily because the fixed cost of the BeadChip is amortized across more samples, not because the chemistry becomes cheaper. A single iScan system can process approximately 1,728 GDA samples per week assuming 24-hour operation with automated array loaders — sufficient throughput for all but the largest national biobanks.

At the time of writing, approximate per-sample costs (consumables only, not including labor, equipment amortization, or bioinformatics) are: $30 to $50 for a GDA-class array (1.8M markers, 8-sample format), $20 to $35 for a GSA v4 (650K markers, 48-sample format), $15 to $25 for a mid-density agricultural array (50K markers, 24-sample format), and $40 to $70 for a custom array of equivalent density, depending on order volume. These figures are estimates that vary with regional pricing, institutional agreements, and batch size, but the relative ordering is stable across markets.

The comparison with sequencing-based genotyping depends on the study's marker requirements. For reference, 0.5× lcWGS with imputation to a large reference panel can recover genotype calls at 5 to 20 million common variants (MAF > 1 percent), far beyond any array, but with variable imputation accuracy — imputation R² values decline with MAF. At approximately $20 to $40 per sample for 0.5× lcWGS, the sequencing option offers more genetic information per dollar when the study can tolerate imputation uncertainty and requires genome-wide coverage beyond what any array provides. The array retains the advantage when: (1) the markers of interest are known, fixed, and well-validated on the array, (2) samples are processed in batches over years and must be directly comparable without re-imputation, (3) the downstream analysis pipeline (PRS, genomic prediction) has been validated specifically on array genotypes and re-validation on imputed sequencing genotypes is costly, or (4) regulatory or translational research requirements demand genotyping at precisely specified markers with auditable QC rather than genome-wide probabilistic genotype calls.

A practical case illustrates the trade-off. A popcorn breeding program submits 1,500 doubled haploid lines for genomic prediction. The program has an existing training population of 300 lines genotyped on the MaizeSNP50 array and phenotyped across three environments. Genotyping the 1,500 new lines on the same array (~$25 per sample) produces genotype calls directly comparable to the training population; the total consumables cost is approximately $37,500. Switching to 0.5× lcWGS would require imputing both the new lines and the training population to a common variant set, re-estimating the genomic prediction model on the imputed genotypes, and validating that prediction accuracy on imputed data matches the accuracy previously demonstrated on array genotypes — a bioinformatic investment of weeks to months whose cost in personnel time may exceed the difference in consumables. Conversely, if the same program were starting from scratch with no existing genotype data and no validated prediction model, lcWGS at 0.5× followed by imputation and model training on the full cohort would likely yield higher prediction accuracy for the same or lower total cost.

The practical recommendation for core facilities: maintain both capabilities — array-based genotyping for ongoing, standardized, high-throughput projects with established marker sets, and sequencing-based genotyping for discovery, non-model species, and projects where marker content requirements are not yet settled. These are complementary, not competing, technologies, and a facility that offers both with informed guidance on platform selection provides a service that neither an array-only nor a sequencing-only provider can match. For researchers planning GWAS on array-genotyped cohorts or transitioning from array-based to sequencing-based genotyping, GWAS analysis services and genotyping by sequencing (GBS) provide complementary paths from genotype data to biological insight.

FAQ

What is the difference between Infinium I and Infinium II assay chemistry?

Infinium I uses two bead types per SNP, one for each allele, and discriminates alleles by single-base extension with hapten-labeled dideoxynucleotides followed by immunohistochemical staining. Infinium II uses a single bead type per SNP with two-color fluorescent staining after extension. Infinium I typically achieves marginally higher call rates; Infinium II enables higher marker density per chip. The Infinium EX chemistry released in 2025 reduces DNA input requirements and processing time for both assay types.

How much DNA do I need for an Infinium array?

The standard protocol requires 200 ng of double-stranded genomic DNA at a minimum concentration of 50 ng/µL. DNA should have a 260/280 ratio between 1.8 and 2.0, assessed by spectrophotometry, but concentration should be measured by fluorometry (Qubit or QuantiFluor) to avoid overestimation from degraded nucleic acid. Genomic DNA integrity should show a predominant band above 10 kb on an agarose gel or fragment analyzer.

What are the most important QC metrics for Infinium genotyping data?

Four metrics matter most: call rate (SNP-level ≥ 98 percent, sample-level ≥ 95 percent), GenCall score (per-genotype confidence; ≥ 0.7 acceptable, < 0.5 should be discarded), cluster separation (per-SNP cluster resolution; ≥ 0.6 for inclusion), and reproducibility (blind duplicate concordance ≥ 99.5 percent). These metrics should be assessed jointly — a SNP that passes each individual threshold marginally may still carry systematic error.

When should I use a custom array instead of a commercial one?

Custom arrays are warranted when no commercial array exists for your species, when existing arrays are designed for genetically distant populations and perform poorly due to SNP ascertainment bias, or when your study requires a specific marker set (from prior GWAS, QTL mapping, or functional annotation) not represented on commercial products. Expect a four-stage cycle: SNP discovery, in silico scoring with Illumina's ADT, manufacturing (8-12 weeks), and pilot validation on 48 to 96 samples.

How does array-based genotyping compare to GBS or lcWGS on cost?

Approximate 2025 prices: GDA-class array (1.8M markers) $30-$50/sample; GSA v4 (650K markers) $20-$35/sample; agricultural array (50K markers) $15-$25/sample; 0.5× lcWGS $20-$40/sample; GBS (10K-100K SNPs) $15-$40/sample. The cost crossover depends on marker requirements, batch size, and whether existing data must remain comparable. Arrays retain an advantage when marker comparability across long time horizons is non-negotiable.

What is cluster compression and how should I handle it?

Cluster compression — when the three genotype clusters are shifted toward the origin of the intensity plot, reducing their separation — is the most common marker-level QC failure. It has three primary causes: secondary SNPs in the probe-binding region reducing hybridization efficiency, paralogous sequences co-hybridizing and adding signal to both channels, and copy number variation producing more than three clusters. SNPs with severely compressed clusters (cluster separation < 0.6) should be excluded. Moderate compression (0.6-0.8) warrants manual re-clustering in GenomeStudio.

Can I combine array genotypes with sequencing-based genotypes in the same analysis?

Yes, but not naively. Array and sequencing genotypes at overlapping markers must be harmonized — strand orientation verified, reference alleles matched, and genotype calling methods calibrated — before merging. Imputation to a common reference panel is the standard harmonization strategy. For cross-platform studies, genotyping a subset of samples on both platforms enables direct calibration of genotype concordance rates.

References:

  1. Illumina, Inc. Infinium Global Diversity Array with Enhanced PGx — Product Data Sheet (Pub. No. M-GL-00543). 2025. https://knowledge.illumina.com/microarray/general/microarray-general-reference_material-list/000005491
  2. Lepoittevin C, Bodénès C, Chancerel E, et al. Single-nucleotide polymorphism discovery and validation in high-density SNP array for genetic analysis in European white oaks. Molecular Ecology Resources. 2015;15(6):1446-1459. https://doi.org/10.1111/1755-0998.12407
  3. Yang L, Zhang Z, Li J, et al. Gsformer: a dual-architecture deep learning framework with CNN-self-attention and sparse-attention for genomic selection. Genetics Selection Evolution. 2026;58:10. https://doi.org/10.1186/s12711-026-01055-8
  4. Crooijmans RPMA, Gonzalez Prendes R, Colli L, et al. IMAGE001: a new livestock multispecies SNP array to characterize genomic variation in European livestock gene bank collections. Animal Genetics. 2025;56(5):e70039. https://doi.org/10.1111/age.70039
  5. Steemers FJ, Gunderson KL. Whole genome genotyping technologies on the BeadArray platform. Biotechnology Journal. 2007;2(1):41-49. https://doi.org/10.1002/biot.200600213

Related Services

For research use only, not intended for clinical diagnosis, treatment, or individual health assessments.

For research purposes only, not intended for clinical diagnosis, treatment, or individual health assessments.
Speak to Our Scientists
What would you like to discuss?
With whom will we be speaking?

* is a required item.

Contact CD Genomics
Terms & Conditions | Privacy Policy | Feedback   Copyright © CD Genomics. All rights reserved.
Top