Genome-Wide Association Studies (GWAS) with GBS: Experimental Design, Sample Size, and Statistical Power

GWAS — Connecting Genotype to Phenotype via Linkage Disequilibrium

Genome-wide association studies solve a problem that has vexed geneticists since the rediscovery of Mendel: which of the thousands to millions of segregating variants in a genome actually cause phenotypic differences? The principle is elegant — if a marker is physically close enough to a causal variant that they are co-inherited across generations (linkage disequilibrium, LD), then allele frequency at that marker will differ between groups of individuals sorted by phenotype. A marker that reliably co-occurs with disease resistance, elevated yield, or drought tolerance has not necessarily identified the causal mutation, but it has identified its chromosomal neighborhood, and that is often enough to be useful.

GWAS in the GBS era. For most of GWAS history, the limiting factor was marker density. Early plant association panels used a few hundred SSR or RFLP markers, which surveyed such a small fraction of the genome that only large-effect loci in regions of extended LD could be detected — flowering time, major disease resistance genes, obvious morphological traits. Complex polygenic traits — yield, drought tolerance, root architecture — were invisible to these sparse marker sets because the probability that any given marker was in LD with any given causal variant was simply too low.

Genotyping-by-sequencing (GBS) changed this calculus. A single GBS run on 200 individuals delivers 10,000 to 100,000 SNPs distributed across the genome at a per-sample cost that makes population-scale genotyping feasible for non-model species. GBS-derived SNPs are not evenly spaced — they cluster around restriction sites — but their density in genic regions is often higher than in intergenic space, particularly with methylation-sensitive enzymes such as PstI that preferentially sample hypomethylated, gene-rich chromatin. Sharma et al. (2024) demonstrated in autotetraploid potato that GBS markers were concentrated in genic regions and outperformed fixed-content SNP arrays for LD estimation, population structure discrimination, and GWAS detection power, identifying 189 unique QTL associations across 16 tuber traits.

For the researcher planning a GWAS, GBS converts the question from "can I afford enough markers?" to "do I have enough samples and phenotypes to exploit the markers I can afford?"

How GBS enables GWAS in non-model species. GBS bridges a critical infrastructure gap. Model species — Arabidopsis, maize, rice — have reference genomes, dense SNP arrays, and community resources that lower the barrier to GWAS. For a non-model fruit tree, forage grass, or aquaculture species with no array and a draft genome at best, GBS provides a direct path to genome-wide marker discovery. A 2025 BMC Genomics study demonstrated that even in less-studied farmed species without a reference genome, ddRAD-based GBS with a mock reference constructed from pooled reads of three representative samples delivered over 15,000 variants — sufficient for GWAS and genomic prediction — with the entire pipeline from DNA to genotype matrix scalable to 384-sample plates. The practical implication is that GWAS is no longer restricted to species with pre-existing genomic infrastructure; any diploid eukaryote with sufficient nucleotide diversity to produce polymorphic GBS tags is a viable GWAS target.

When GBS is not enough. GBS provides 10,000 to 100,000 SNPs, but these are concentrated in genomic regions adjacent to restriction sites. Traits controlled by variants in genomic deserts — regions with few restriction sites — may be invisible to GBS-based GWAS. For studies where GBS fails to detect expected associations despite adequate sample size and phenotype quality, low-coverage whole-genome sequencing (lcWGS) with imputation to a reference panel, or targeted SNP panels designed from whole-genome resequencing of a subset of individuals, can fill the gap. For a broader overview of genotyping technology options from GBS through ddRAD-seq to whole-genome approaches, see CD Genomics' Genotyping and Genetic Diversity Services overview.

The GWAS-to-targeted-genotyping pipeline. A productive workflow that has become standard in plant breeding programs proceeds in three stages: (1) GBS-based SNP discovery and GWAS in a diversity panel of 200 to 500 individuals, identifying 20 to 100 marker-trait associations at a relaxed significance threshold; (2) conversion of the most promising SNPs to cost-effective targeted assays (KASP or targeted amplicon sequencing) for validation in independent populations — TaqMan SNP genotyping provides high-accuracy validation for panels of 1 to 50 candidate SNPs; and (3) deployment of validated markers in breeding populations for marker-assisted selection or genomic prediction.

Figure 1: GWAS Principle — From GBS-Derived SNPs to Phenotype Association via Linkage Disequilibrium Figure 1: GWAS Principle — From GBS-Derived SNPs to Phenotype Association via Linkage Disequilibrium

Experimental Design — Populations, Sample Size, and Statistical Power

The quality of a GWAS is determined more by experimental design decisions made before a single library is prepared than by which statistical model is applied to the data afterward. Three design choices — population composition, sample size, and phenotyping strategy — collectively determine whether biologically real associations will be detected at genome-wide significance or lost in the noise. A well-designed GWAS on 300 carefully phenotyped individuals genotyped at 30,000 GBS SNPs will reliably outperform a poorly designed study on 500 individuals with noisy phenotypes and 100,000 SNPs. This is not because statistics are unimportant — they are — but because no statistical model can recover signal that was never captured in the first place.

Population types and their trade-offs. Diversity panels — collections of accessions representing the breadth of a species' geographic, phenotypic, and genetic variation — are the most common GWAS population in plants. The advantage is high allelic richness and the ability to survey many historical recombination events. The disadvantage is population structure: allele frequency differences driven by demographic history rather than phenotype can produce spurious associations that persist even after correction. The 2025 sugar beet GWAS by the BMC Plant Biology consortium, which genotyped 94 accessions from 16 countries using 4,609 GBS-derived SNPs, illustrates both the strength and limitation of diversity panels — 35 significant marker-trait associations and 25 candidate genes were identified for root and quality traits, but the relatively modest sample size limited power to detect small-effect loci, a trade-off the authors explicitly acknowledged.

Bi-parental mapping populations (F2, recombinant inbred lines, doubled haploids) eliminate population structure as a concern — all individuals share the same two parental genomes, and any allele frequency divergence between phenotypic extremes must be driven by the trait, not by demographic history. The cost is reduced allelic diversity: only variants segregating in the two parents can be mapped. Multi-parent advanced generation intercross (MAGIC) and nested association mapping (NAM) populations occupy the middle ground, combining the controlled structure of bi-parental crosses with the allelic diversity of multiple founders. Kitony et al. (2026) demonstrated in a rice NAM population with 1,818 recombinant inbred lines across 14 families that genomic prediction accuracy saturated at approximately 500 lines with moderate marker density, while GWAS resolution continued to improve with additional markers, highlighting why the same population may be adequate for one goal and underpowered for another.

Sample size and statistical power. Power in GWAS is a function of four parameters: the phenotypic variance explained by the locus (PVE, or effect size), the minor allele frequency (MAF) at the marker, the significance threshold applied, and the sample size. The relationship is nonlinear and unforgiving — detecting a locus explaining 5 percent of phenotypic variance at genome-wide significance (typically p < 1 × 10⁻⁵ to 5 × 10⁻⁸ for GBS-scale marker sets) requires a minimum of 300 to 500 individuals for traits of moderate heritability (h² ≈ 0.4–0.6), while loci with PVE below 2 percent routinely require sample sizes exceeding 1,000 that are beyond the reach of most single-investigator studies. The 2025 sugar beet GWAS consortium acknowledged this explicitly: their panel of 94 accessions was sufficient to detect the 35 marker-trait associations they reported, but the study was underpowered for the smaller-effect loci that likely contribute to root quality — a pragmatic admission that reflects the reality of working with germplasm collections where the number of available accessions, not the budget for sequencing, is the binding constraint.

A practical power analysis framework uses the effective number of independent markers (Me) rather than the raw SNP count when calculating Bonferroni thresholds, because GBS SNPs in LD are not independent tests. The SimpleM method estimates Me from the eigenvalues of the LD matrix, and the significance threshold becomes α / Me (where α = 0.05), which is typically orders of magnitude less conservative than α divided by the raw SNP count. For a typical GBS dataset with 30,000 SNPs but an Me of 5,000 to 8,000, the Bonferroni threshold softens to approximately p < 6 × 10⁻⁶ to 1 × 10⁻⁵, within reach of moderately powered studies. GCTA and the R package genpwr provide tools for formal power estimation given user-specified PVE, MAF, and sample size. A useful pre-study exercise is to simulate: given your expected sample size, what is the smallest PVE you can detect at 80 percent power? If that PVE is larger than the effect sizes reported in comparable published studies for your trait, the study may be underpowered and should either expand sample size or refocus on traits with simpler genetic architecture.

Common design pitfalls. Three recurring mistakes compromise otherwise well-executed GWAS. First, unbalanced sampling — collecting phenotypes from a core set of genotypes but genotyping a larger set that includes many individuals without phenotypes — wastes sequencing budget on samples that contribute nothing to association tests. Second, ignoring environmental heterogeneity within a single field site — soil gradients, edge effects, and irrigation patterns that differentially affect plots within the same trial — introduces noise that no amount of genotyping can overcome. Third, using a single-year phenotype for traits with substantial genotype-by-year interaction, which can produce associations that fail to replicate because they capture year-specific rather than general genetic effects. The fix for all three is straightforward: phenotype all genotyped individuals, replicate measurements within and across environments, and analyze multi-year data before declaring novel associations.

Figure 2: GWAS Experimental Design Matrix — Population Type vs. Sample Size vs. Power Figure 2: GWAS Experimental Design Matrix — Population Type vs. Sample Size vs. Power

Phenotyping — The Determinant of GWAS Success

The statistical machinery of GWAS is indifferent to the biological meaning of the numbers it receives, which makes phenotyping the single most consequential step in the pipeline. GWAS cannot rescue phenotypes that are noisy, unreplicated, or confounded.

Heritability and measurement precision. Broad-sense heritability (H²) on an entry-mean basis sets the upper bound on what any GWAS can detect: if H² is 0.3, the best-powered study in the world cannot explain more than 30 percent of the observed phenotypic variance. Heritability is improved by replication — multiple plants per plot, multiple plots per environment, multiple environments per genotype. For field-based traits with moderate to low heritability (H² < 0.4), a minimum of two to three replicate plots per genotype, planted in a randomized complete block design, is recommended. A common mistake in first-time GWAS designs is investing in genotyping density at the expense of phenotyping replication — 100,000 SNPs on 300 unreplicated field plots will produce fewer reliable associations than 10,000 SNPs on 300 genotypes replicated across three environments with two blocks each.

Multi-environment trials and genotype-by-environment interaction. Genotype-by-environment interaction (G×E) — the phenomenon where a genotype that performs well in one location or year performs differently in another — is pervasive for complex traits and can completely obscure genetic signals when phenotypes are collected from a single environment. Best linear unbiased estimates (BLUEs) across environments provide the most reliable single-value phenotype for GWAS when G×E is modest. When G×E is strong and trait expression differs qualitatively across environments, environment-specific GWAS — running separate analyses for each environment and intersecting the results — identifies loci that are stable versus environment-dependent, an approach particularly valuable for drought tolerance and disease resistance where environment-specific adaptation is the trait of interest, not noise.

Measuring phenotypes on the same individuals used for GBS may seem obvious, but the logistical reality — coordinating field planting, tissue collection for DNA extraction, and trait measurement across the growing season — requires planning that begins months before the first seed is sown. A 2025 tropical maize GWAS that compared GBS-derived and array-derived SNPs across well-watered and water-stressed conditions found that combining genotyping approaches increased GWAS resolution, but the study's most important finding was that water regime-specific GWAS detected loci invisible in the combined analysis, underscoring that environmental context is not a statistical nuisance but a biological signal.

Modern phenotyping technologies — drone-based multispectral imaging, automated weighing and imaging platforms, handheld near-infrared spectrometers — are increasingly integrated with GWAS pipelines, enabling measurement of traits (canopy temperature, growth rate, chlorophyll content) at temporal resolution and throughput that manual measurement cannot match. However, these technologies measure secondary phenotypes correlated with the trait of interest, not the trait itself, and the genetic architecture of a secondary phenotype may differ from that of the trait it proxies. High-throughput phenotyping expands the number of traits that can be studied but does not eliminate the need for careful validation that the measured phenotype is biologically meaningful for the question at hand.

Figure 3: Phenotyping Design — Replication, Multi-Environment Trials, and Heritability Partitioning Figure 3: Phenotyping Design — Replication, Multi-Environment Trials, and Heritability Partitioning

GBS Data Processing for GWAS — From FASTQ to Filtered Genotype Matrix

The bioinformatic pipeline that converts raw GBS reads into an analysis-ready genotype matrix requires methodical attention to filtering, imputation, and format conversion. Shortcuts at this stage produce downstream artifacts that are indistinguishable from genuine biological signals.

SNP discovery and genotype calling. Two pipelines dominate GBS data processing for GWAS. The TASSEL-GBS pipeline (version 5) uses a reference-genome-guided approach: reads are aligned with Bowtie2 or BWA, SNPs are called from tag-level alignments, and genotypes are exported in VCF or HapMap format. TASSEL has been validated at production scale — large-scale cereal breeding programs have processed tens of thousands of breeding lines through TASSEL-GBS v5, applying Fisher's exact test and chi-square filters for SNP quality alongside MAF and call-rate thresholds. The alternative Stacks pipeline (denovo_map.pl or ref_map.pl) assembles loci de novo from the GBS reads or against a reference, calls variants with gstacks, and exports filtered genotype matrices through the populations module. Stacks excels for non-model species without a reference genome, where its paired-end contig assembly produces locus sequences that can be used for downstream primer design.

Imputation. GBS datasets typically contain 20 to 40 percent missing data because restriction sites that are polymorphic in some individuals (producing a sequenced tag) are absent in others (producing missing data at that locus). Imputation fills these gaps using LD information from neighboring markers. Beagle (version 5) uses a haplotype-clustering model that scales efficiently to thousands of samples and hundreds of thousands of markers. LD-kNNi, implemented in TASSEL, imputes missing genotypes by finding the k nearest neighbors in LD space, which is computationally faster than Beagle for smaller datasets but may underperform when LD is low or marker density is sparse. The choice matters: imputation quality directly affects GWAS power because poorly imputed genotypes add noise to association tests, and systematic differences in imputation accuracy between rare and common alleles can bias effect-size estimates. A practical rule is to impute only after filtering out SNPs with initial call rates below 50 percent, and to apply a post-imputation accuracy filter — removing SNPs with imputation quality scores (DR² or Rsq from Beagle) below 0.6 to 0.8 before proceeding to GWAS.

Filtering cascade. Standard pre-GWAS filters applied sequentially are: (1) individual call rate ≥ 80 percent (remove samples with excessive missing data, often indicating poor DNA quality or library failure), (2) SNP call rate ≥ 70 to 80 percent (remove loci recovered in too few samples to be informative), (3) minor allele frequency ≥ 5 percent (rare variants lack power in moderate-sized panels and are enriched for sequencing errors), and (4) Hardy-Weinberg equilibrium p > 0.001 (flag potential genotyping errors, though genuine biological departures from HWE exist in structured populations and should not be blindly excluded). LD pruning — retaining one SNP per LD block (r² < 0.8 within a sliding window) — reduces the marker set to approximately independent tests and is essential before methods that assume marker independence (PCA, STRUCTURE).

Reference genome considerations. The availability of a reference genome — even a fragmented draft — transforms GBS data processing. With a reference, reads are aligned with BWA-MEM or Bowtie2 and variants called with GATK or freeBayes, with the substantial advantage that SNP positions are anchored to chromosomes and comparable across studies. Without a reference, Stacks' de novo pipeline assembles reads into tag-level loci, but these loci are defined by sequence rather than genomic position — a Tag 42 in one study bears no relationship to a Tag 42 in another, making cross-study meta-analysis impossible. The 2025 BMC Genomics fine-tuning study found that a mock reference constructed from pooled GBS reads of just three representative samples within a species performed nearly as well as a true reference genome for SNP calling, an approach that has become standard for non-model species where a full reference is unavailable but the study demands positional information for candidate gene searches. For species where genomic resources are intermediate — a transcriptome assembly exists but no genome — transcriptome-anchored GBS, in which reads are aligned to the transcriptome rather than the genome, biases marker recovery toward expressed genes, which may be advantageous for trait mapping but systematically misses regulatory variants in non-genic regions.

Figure 4: GBS Data Processing Pipeline — From Raw Reads to Filtered Genotype Matrix Figure 4: GBS Data Processing Pipeline — From Raw Reads to Filtered Genotype Matrix

Statistical Models — From GLM to BLINK

The evolution of GWAS statistical models over the past two decades tells a story of progressive sophistication in controlling false positives without sacrificing true discovery power. Understanding this evolution matters because model choice directly affects which associations are reported and which are missed.

The GLM-to-MLM progression. A naive general linear model (GLM) regresses phenotype on genotype, one SNP at a time: y = μ + SNP + ε. In structured populations — where certain alleles are common in one subpopulation and rare in another for reasons unrelated to the trait — GLM produces genomic inflation factors (λ) well above 1.0 and QQ plots that deviate from the diagonal early, indicating pervasive false positives. The mixed linear model (MLM) adds a random polygenic effect whose covariance structure is the kinship matrix (K): y = μ + SNP + u + ε, where u ~ N(0, Kσ²_g). The kinship matrix — estimated from the genotype data itself as identity-by-state (IBS) or allele-sharing coefficients — absorbs the confounding effects of population structure and cryptic relatedness. Adding the first 5 to 10 principal components (PCs) as fixed-effect covariates alongside the kinship random effect (the Q+K model) further controls residual stratification. GAPIT, GEMMA, and EMMAX are the three most widely used implementations, with GAPIT offering the broadest menu of models through a unified R interface.

The FarmCPU-to-BLINK leap. Compressed MLM (CMLM) and enriched CMLM (ECMLM) improved computational efficiency by clustering individuals into groups, but the real methodological breakthrough came with multi-locus iterative models. FarmCPU (Fixed and random model Circulating Probability Unification) alternates between a fixed-effect model that tests markers one at a time using pseudo-quantitative trait nucleotides (pseudo-QTNs) selected from a previous iteration as covariates, and a random-effect model that re-estimates these pseudo-QTNs. BLINK (Bayesian-information and Linkage-disequilibrium Iteratively Nested Keyway) eliminates the random-effect step entirely, using BIC-based model selection with explicit LD pruning (r² > 0.7) to build the set of covariate markers. Fatima (2025) compared eight GWAS models across heritability (0.3–0.8) and polygenicity (50 vs. 100 QTLs) gradients in simulated plant trait data, confirming that BLINK consistently detects the highest number of true positives, particularly for moderately heritable traits, while MLMM provides superior mapping resolution closer to the causal variant.

Multiple testing correction. The Bonferroni threshold — α divided by the number of tests — is the most conservative approach and the most commonly reported. Using the effective number of independent markers (Me from SimpleM) rather than the raw SNP count makes Bonferroni practical for GBS-scale data. The Benjamini-Hochberg false discovery rate (FDR) at q < 0.05 is less conservative and appropriate when the goal is candidate gene discovery for downstream validation rather than definitive causal claims. Regardless of the threshold chosen, the QQ plot — which compares observed versus expected p-value distributions — should show alignment with the diagonal for the vast majority of SNPs, with deviation only in the extreme tail where true associations reside. A λ value near 1.0 (typically < 1.05 for plant GWAS) indicates adequate structure control.

Model selection in practice. The 2025 Bio-Protocol paper by the GAPIT development team recommends a staged approach: begin with a computationally efficient model (FarmCPU or BLINK) for the initial scan, then validate the top associations with a traditional Q+K MLM. A systematic comparison across eight models in simulated plant datasets with varying heritability (0.3–0.8) and polygenicity (50 vs. 100 QTLs) confirmed that BLINK detects the most true positives, MLMM achieves the finest mapping resolution, and FarmCPU provides the best balance for large datasets where computational cost matters. The key operational rule: if BLINK and FarmCPU report similar sets of significant SNPs, confidence in those associations is high; if the two models diverge substantially, population structure or model misspecification may be at play, and a more conservative MLM analysis is warranted before committing resources to validation genotyping.

Manhattan plots — genomic position on the x-axis, −log₁₀(p) on the y-axis — remain the standard visualization. Peaks rising above the significance threshold line, particularly those spanning multiple consecutive SNPs in LD, are the primary output of GWAS and the starting point for biological interpretation. The next step — determining which gene(s) underlie a peak — requires a different set of tools entirely.

Figure 5: GWAS Statistical Models — GLM vs. MLM vs. FarmCPU vs. BLINK Comparison Figure 5: GWAS Statistical Models — GLM vs. MLM vs. FarmCPU vs. BLINK Comparison

From GWAS Peak to Candidate Gene — Validation and Functional Follow-Up

A significant GWAS peak is a hypothesis, not a conclusion. The genomic interval defined by LD decay around the lead SNP typically contains dozens to hundreds of genes in species with large intergenic regions, and narrowing this list to a tractable set of candidates requires integrating multiple lines of evidence.

LD decay and candidate interval. Linkage disequilibrium decays with physical distance at a rate that varies enormously across species — maize, an outcrosser with a large effective population size, has LD decay to r² < 0.2 within 1 to 2 kb, while soybean, a selfer, maintains LD blocks extending 100 to 150 kb. The candidate interval — defined as the genomic region where r² between the lead SNP and surrounding markers exceeds 0.4 to 0.6 — determines the number of genes that must be evaluated. In maize, a GWAS peak typically implicates fewer than 5 genes; in wheat or soybean, the same peak may span 50 to 200 genes. Species-specific LD decay should be estimated empirically from the study's own genotype data using PopLDdecay or the --r2 command in PLINK, rather than relying on literature values that may not reflect the specific populations and marker sets in use.

Candidate gene annotation and prioritization. The minimal annotation pipeline maps lead SNPs and their LD proxies to the reference genome, extracts all genes within the candidate interval, and queries functional databases (Gene Ontology, KEGG, InterPro, Pfam) for terms relevant to the trait. Genes with expression in the trait-relevant tissue (root for drought tolerance, flower for flowering time, seed for grain quality) are prioritized. When RNA-seq data exist for the same or related populations, expression QTL (eQTL) analysis — testing whether the lead GWAS SNP is also associated with expression level of nearby genes — provides a direct functional link: a SNP associated with both disease resistance and the expression of an NBS-LRR gene in leaf tissue is a far stronger candidate than a SNP associated only with the phenotype.

Beyond standard GWAS — advanced approaches for complex traits. For traits where standard SNP-based GWAS yields few or no significant associations despite adequate sample size, several methodological extensions can recover signal. Transcriptome-wide association studies (TWAS) integrate gene expression data with genotype data, testing for associations between predicted expression levels and phenotype. A landmark 2024 TWAS in soybean used RNA-seq from 622 accessions to identify 29,286 genes with genetically regulated expression, then associated these expression levels with multiple agronomic traits — discovering a novel pod-color gene (L2) that standard GWAS on the same population had missed because the causal variant was a structural rearrangement, not a SNP captured by the GBS or array markers. Multi-omics integration — combining GWAS with metabolomics, proteomics, or epigenomic data — is an emerging frontier. For researchers interested in GWAS analysis services, pipelines that integrate multiple statistical models, imputation methods, and candidate gene annotation provide a systematic framework for converting GBS data into biological insight.

QTL-seq and BSA-seq for independent validation. Bulked segregant analysis with sequencing (BSA-seq or QTL-seq) provides a rapid, cost-effective validation approach that is independent of the GWAS framework. The method crosses two individuals with extreme phenotypes, pools DNA from the phenotypic extremes in the segregating progeny, sequences the pools, and identifies genomic regions where allele frequencies diverge between the high and low pools (quantified as the ΔSNP-index or G-statistic). Because BSA-seq uses a bi-parental population rather than a diversity panel, it probes a different recombination landscape — a QTL detected by both GWAS and BSA-seq has survived two independent tests in genetically distinct populations, making it a high-confidence candidate for marker development. A 2024 peanut pod shell thickness study used BSA-seq with four statistical algorithms (ΔSNP-index, Euclidean distance, G-value, Fisher's exact test) to identify two major QTLs explaining 31 to 32 percent and 16 to 17 percent of phenotypic variance respectively, and converted the top markers to KASP assays for breeding deployment — a workflow that mirrors the path from GWAS discovery to breeding application.

Practical deployment — validated markers in breeding programs. The endpoint of a GWAS pipeline is not a publication but a set of markers that a breeder can use to make selection decisions. Converting GWAS hits to breeder-friendly markers involves three practical steps. First, the lead SNPs are converted to KASP or TaqMan assays — single-plex, qPCR-based genotyping reactions that can be run on standard laboratory equipment without the bioinformatic infrastructure that GBS requires. For panels of 1 to 50 validated SNPs, KASP is the most economical format and can be deployed at scale: a single technician with a qPCR instrument can process hundreds of samples per day for a 10-SNP panel. Second, the assays are validated on an independent population — ideally one that shares genetic background with the breeding material but was not part of the original GWAS discovery panel. Markers that fail to associate with the trait in the validation population are discarded regardless of their GWAS p-value. Third, the surviving markers are incorporated into the breeding program's decision workflow: either as fixed selection criteria (must carry the resistance allele at markers X, Y, and Z) or as weighted components of a genomic selection index where validated GWAS hits receive higher weight than anonymous GBS SNPs. GWAS analysis services that include candidate marker conversion to KASP or TaqMan format provide a direct pipeline from discovery to deployment.

An end-to-end GWAS project — from the initial decision to genotype a diversity panel through to validated markers deployed in a breeding program — is a substantial undertaking that typically spans 18 to 36 months. The timeline is dominated not by sequencing or bioinformatics but by phenotyping: a single growing season for annual crops, potentially multiple seasons for perennials, and additional time for multi-environment validation. Sequencing and analysis — GBS library preparation through to the first Manhattan plot — can be completed in 8 to 12 weeks once DNA is extracted. Researchers who budget adequate time for phenotyping and allocate 10 to 15 percent of the total project budget to pilot genotyping of 16 to 24 samples ahead of the full cohort are rewarded with datasets that yield interpretable, replicable results rather than frustratingly noisy Manhattan plots.

Figure 6: GWAS-to-Candidate-Gene Pipeline — LD Decay, Annotation, and QTL-seq Validation Figure 6: GWAS-to-Candidate-Gene Pipeline — LD Decay, Annotation, and QTL-seq Validation

FAQ

What is a genome-wide association study (GWAS)?

GWAS tests statistical associations between genetic markers (typically SNPs) distributed across the genome and a phenotype of interest, exploiting linkage disequilibrium — the non-random association of alleles at nearby loci — to identify chromosomal regions harboring causal variants. A significant association does not identify the causal mutation directly but pinpoints its genomic neighborhood for functional follow-up.

How many samples do I need for a GBS-based GWAS?

A minimum of 200 to 300 individuals for detecting loci explaining 5 percent or more of phenotypic variance in traits with moderate heritability. Loci with smaller effects or lower allele frequencies require 500 to 1,000 or more individuals. The effective number of independent markers (Me), trait heritability, and desired power all influence the required sample size, and tools such as GCTA and genpwr enable formal power calculations before committing to a study design.

What is the difference between GWAS and QTL mapping?

GWAS uses natural populations (diversity panels) and exploits historical recombination accumulated over many generations, providing higher mapping resolution but requiring correction for population structure. QTL mapping uses bi-parental experimental crosses and follows recombination events occurring in a single generation, yielding lower resolution but higher statistical power per marker and freedom from population structure artifacts.

Which statistical model should I use for GWAS?

For most plant GWAS with moderate population structure, a mixed linear model incorporating kinship and principal components (Q+K MLM), as implemented in GAPIT, GEMMA, or EMMAX, provides adequate false-positive control. For studies with complex population structure or when maximizing power is critical, BLINK or FarmCPU — both multi-locus iterative models — offer superior detection power while maintaining false-positive control.

Do I need a reference genome for GWAS?

A reference genome substantially improves GWAS — it enables candidate gene annotation, LD decay estimation in physical coordinates, and cross-study comparison of associated loci. However, GBS-based GWAS can be conducted de novo using Stacks to assemble tag-level loci; associations are reported as tag sequences rather than chromosomal positions, and significant tags can be BLAST-searched against related species' genomes for tentative annotation.

What is the difference between GLM and MLM in GWAS?

A general linear model (GLM) tests marker-phenotype associations without accounting for relatedness among individuals, producing inflated false-positive rates in structured populations. A mixed linear model (MLM) includes a random polygenic effect whose covariance is the kinship matrix, which absorbs confounding due to population structure and cryptic relatedness. The Q+K variant adds principal components as fixed-effect covariates for additional structure control.

How is statistical significance determined in GWAS?

The Bonferroni correction divides the significance threshold α (typically 0.05) by the number of independent tests. Using the effective number of independent markers (Me) rather than the raw SNP count reduces overcorrection from LD. The Benjamini-Hochberg false discovery rate (FDR < 0.05) is a less conservative alternative appropriate for candidate gene discovery. Manhattan plots display −log₁₀(p) for each SNP against genomic position, with significant SNPs appearing as peaks above the threshold line.

What comes after a GWAS identifies significant SNPs?

Significant SNPs are validated in an independent population using targeted genotyping (KASP or amplicon sequencing). The candidate genomic interval is defined by local LD decay around the lead SNP. Genes within the interval are annotated and prioritized by functional relevance to the trait. QTL-seq (BSA-seq) in bi-parental populations provides orthogonal validation that the locus is detected in a genetically distinct background. Validated markers proceed to marker-assisted selection or are incorporated into genomic prediction models.

References:

  1. Zuo Z, Li M, Liu D, Li Q, Huang B, Ye G, Wang J, Tang Y, Zhang Z. GWAS procedures for gene mapping in diverse populations with complex structures. Bio-Protocol. 2025;15(8):e5284. https://doi.org/10.21769/BioProtoc.5284
  2. Li D, Wang Q, Tian Y, Lyu X, Zhang H, Hong H, Gao H, Li YF, Zhao C, Wang J, Wang R, Yang J, Liu B, Schnable PS, Schnable JC, Li YH, Qiu LJ. TWAS facilitates gene-scale trait genetic dissection through gene expression, structural variations, and alternative splicing in soybean. Plant Communications. 2024;5(10):101010. https://doi.org/10.1016/j.xplc.2024.101010
  3. de Pontes FCF, Machado IP, Silveira MVDS, Lobo ALA, Sabadin F, Fritsche-Neto R, DoVale JC. Combining genotyping approaches improves resolution for association mapping: a case study in tropical maize under water stress conditions. Frontiers in Plant Science. 2025;15:1442008. https://doi.org/10.3389/fpls.2024.1442008
  4. Bahjat NM, Yildiz M, Nadeem MA, Morales A, Wohlfeiler J, Baloch FS, Tunçtürk M, Koçak M, Chung YS, Grzebelus D, Sadik G, Kuzğun C, Cavagnaro PF. Population structure, genetic diversity, and GWAS analyses with GBS-derived SNPs and silicoDArT markers unveil genetic potential for breeding and candidate genes for agronomic and root quality traits in an international sugar beet germplasm collection. BMC Plant Biology. 2025;25:523. https://doi.org/10.1186/s12870-025-06525-7
  5. Fischer D, Tapio M, Bitz O, Iso-Touru T, Kause A, Tapio I. Fine-tuning GBS data with comparison of reference and mock genome approaches for advancing genomic selection in less studied farmed species. BMC Genomics. 2025;26:111. https://doi.org/10.1186/s12864-025-11296-4
  6. Clauw P, Ellis TJ, Liu HJ, Sasaki E. Beyond the standard GWAS — a guide for plant biologists. Plant and Cell Physiology. 2025;66(4):431-443. https://doi.org/10.1093/pcp/pcae079
  7. Liu H, Zheng Z, Sun Z, Qi F, Wang J, Wang M, Dong W, Cui K, Zhao M, Wang X, Zhang M, Wu X, Wu Y, Luo D, Huang B, Zhang Z, Cao G, Zhang X. Identification of two major QTLs for pod shell thickness in peanut (Arachis hypogaea L.) using BSA-seq analysis. BMC Genomics. 2024;25:101. https://doi.org/10.1186/s12864-024-10005-x
  8. Sharma SK, McLean K, Hedley PE, Dale F, Daniels S, Bryan GJ. Genotyping-by-sequencing targets genic regions and improves resolution of genome-wide association studies in autotetraploid potato. Theoretical and Applied Genetics. 2024;137:180. https://doi.org/10.1007/s00122-024-04651-8
  9. Kitony JK, Reyes VP, Sunohara H, Tasaki M, Yamasaki M, Mori J, Shimazu A, Nishiuchi S, Michael TP, Doi K. Efficient genomic prediction at reduced training size and moderate marker density in an expanded aus-NAM population of rice. bioRxiv. 2026. https://www.biorxiv.org/content/10.64898/2026.04.28.721500v1
  10. Fatima N. Evaluating GWAS model performance across heritability and polygenicity gradients in simulated plant trait data. Preprints.org. 2025. https://www.preprints.org/manuscript/202503.2269

For research use only, not intended for clinical diagnosis, treatment, or individual health assessments.

For research purposes only, not intended for clinical diagnosis, treatment, or individual health assessments.
Speak to Our Scientists
What would you like to discuss?
With whom will we be speaking?

* is a required item.

Contact CD Genomics
Terms & Conditions | Privacy Policy | Feedback   Copyright © CD Genomics. All rights reserved.
Top