Training Population Design for Genomic Selection: What Determines Prediction Accuracy?
Genomic selection (GS) has transformed crop and livestock breeding by predicting genomic estimated breeding values (GEBV) for unphenotyped selection candidates using genome-wide marker profiles. While quantitative geneticists frequently compare GBLUP, Bayesian models, and machine-learning approaches, training-population composition, relatedness, and phenotype quality can influence prediction accuracy as much as, or more than, differences among commonly used prediction models in many breeding datasets. A poorly matched training population can limit prediction even when marker density is high. This guide focuses on the practical question of who should enter the training population, how to balance relatedness and diversity, how to recognize when a training set is no longer representative, and how to validate an operational training set before selection decisions are made.
Key takeaways
- Genetic relatedness between the training population and target selection candidates is a major driver of prediction accuracy, but it should be optimized jointly with training size and genetic diversity.
- Phenotype reliability can place a strong ceiling on genomic prediction; multi-environment testing, appropriate replication, and field-error adjustment are often as important as adding more genotyped individuals.
- Training-population size shows diminishing returns, and optimization methods such as CDmean or PEVmean can improve sampling efficiency in some breeding populations when phenotyping capacity is limited.
- Training sets should be reviewed as breeding populations change, with updates triggered by declining forward-prediction performance, new parents, or reduced genetic connectivity rather than by a fixed number of cycles.
- Marker density should be sufficient for the linkage disequilibrium structure of the target population, but denser genotyping cannot compensate for poorly matched phenotypes or genetically disconnected training material.
Core determinants of genomic prediction accuracy
Who should be included in a genomic selection training population? A useful training population should represent the genetic backgrounds present in the target selection candidates, contain reliable phenotypic information for the traits of interest, and reflect the environments in which future candidates will be evaluated. The best design is therefore population- and trait-specific rather than a fixed sample-size recipe.
Quantitative genetics theory often approximates expected genomic prediction accuracy (rg,ĝ) as a function of training population size (NTP), narrow-sense trait heritability (h2), and the effective number of independent chromosome segments (Me = 2 × Ne × L / ln(4 × Ne × L)):
Expected prediction accuracy ≈ √[(NTP × h2) / (NTP × h2 + Me)]
This theoretical approximation illustrates why larger training populations and more reliable phenotypes can improve expected accuracy, while genetically complex populations with more independent chromosome segments can require more information. Realized accuracy can differ substantially because of training-target relatedness, marker coverage, trait architecture, population structure, genotype-by-environment interaction, and the validation design used to estimate performance. For an overarching perspective on quantitative models and breeding pipelines, see the internal primer on genomic selection in plant and animal breeding.
Genetic relatedness versus training set size
Genetic relatedness between the training population and the target candidate set is consistently important because closely related individuals are more likely to share linkage disequilibrium phase and recent haplotypes. However, relatedness and sample size should not be treated as competing variables with a universal winner. In winter wheat, Edwards et al. (2019) showed that prediction accuracy generally increased with training-set size and that related crosses in training and validation sets usually produced higher accuracy than unrelated crosses. Their results support a practical rule: at a comparable training size, genetically connected material is often more informative, but a sufficiently large and diverse training population can partly compensate for lower relatedness. The relevant question is therefore whether additional samples add new information about the target population rather than simply increasing the row count.
Phenotyping quality and multi-environment trials
Genomic prediction models cannot recover signal that is not present in the phenotype data. Field heterogeneity, unbalanced historical trials, low replication, inconsistent trait definitions, or poorly connected environments can reduce phenotype reliability and limit predictive ability. Multi-environment trials (MET), appropriate spatial or mixed-model adjustments, and well-defined entry means or breeding-value estimates can help produce more informative training phenotypes. Bazzer et al. (2025) showed in oat that incorporating historical multi-environment breeding data increased genomic prediction accuracy compared with using data from a single year. Spatial adjustment remains a project-specific analytical choice and should be evaluated from the field design and residual structure rather than assumed to improve every dataset.
Training population design across breeding scenarios
The optimal strategy for assembling a training set depends on program maturity, germplasm structure, trait architecture, and the intended deployment population. The population sizes below are illustrative planning ranges rather than universal minimum requirements. Appropriate training size should be validated for the actual trait, candidate population, phenotype reliability, mating design, and prediction objective.
For research groups establishing initial marker platforms and determining whether arrays, GBS, or low-pass sequencing best suit their cohort, consult the practical comparison on choosing between LC-WGS, WGS, GBS, and SNP arrays for breeding.
| Breeding program scenario | Primary genetic objective | Illustrative population composition | Key design challenge | Optimization focus |
|---|---|---|---|---|
| New GS Program Launch | Establish baseline prediction models | Historical advanced lines plus primary elite parents; often several hundred well-phenotyped entries | Unbalanced historical phenotype data across years | Connect historical trials, improve phenotype reliability, and minimize marker incompatibility |
| Closed Bi-parental / Multi-family | Predict progeny within defined crossing blocks | Parents, full-sibs, and half-sibs represented across the active families | Differences in LD phase and genetic background across families | Family-aware validation and adequate representation of each target family |
| Multi-Parent Population (MAGIC / NAM) | Predict recombinant lines and evaluate complex traits | Representative recombinant lines anchored to known founders | Unequal founder representation and segregation distortion | Maintain founder coverage and evaluate optimized training subsets |
| Continuous Commercial Pipeline | Sustain prediction across successive selection cycles | Rolling training archive combining informative historical data with newly phenotyped advanced lines | Declining relatedness, new parents, and changes in target environments | Monitor forward prediction and refresh when representativeness declines |
Multi-parent populations and elite diversity panels
Structured multi-parent populations can be especially useful for studying training-set design because their founder contributions are defined. Flores-Saavedra et al. (2026) applied genomic prediction in an eggplant MAGIC population and used fitted models to predict 141 lines that had not been phenotyped for the evaluated water-stress traits. Predictive accuracy varied substantially among traits, illustrating an important limitation: the presence of a structured population does not guarantee uniformly high prediction performance. In spinach, Bhattarai et al. (2025) found that GWAS-informed marker sets could improve prediction in some validation settings, but across-population prediction remained modest. This contrast reinforces the need to test whether a training design transfers to the genetic background in which the model will actually be used.
Algorithmic optimization of training set composition
When phenotyping capacity is limited, the problem is not simply how many lines can be genotyped, but which lines should receive the most expensive and informative phenotyping. Random sampling can over-represent large sibling groups or genetic clusters while leaving other parts of the candidate space poorly represented. Training-population optimization methods use genomic relationships between candidate and training individuals to identify subsets that are expected to provide more useful predictive information.
Coefficient of Determination (CDmean) and Prediction Error Variance (PEV)
Common optimization criteria include:
- CDmean: Seeks to maximize the average coefficient of determination, or expected reliability, of predicted breeding values for the target candidate set.
- PEVmean: Seeks to minimize the mean prediction error variance across target candidates using their genomic relationships with the proposed training set.
- Stratified Clustering Sampling: Uses PCA, clustering, pedigree groups, or heterotic pools to reduce redundant sampling and retain representation of distinct genetic backgrounds.
Berro et al. (2019) evaluated training-population optimization strategies for genomic selection and showed that optimized subsets can improve predictive efficiency relative to random selection in some scenarios, although the advantage depends on the target population, trait, training size, and optimization criterion. The practical benefit is greatest when only a subset of a much larger nursery can be phenotyped. Optimization should therefore be judged by realistic cross-validation against the intended candidate set rather than assumed to outperform random sampling in every breeding program. For end-to-end quantitative data processing and algorithmic selection support, researchers can leverage specialized agricultural genomic data analysis services.
When is the training population no longer representative?
A training set can remain large yet lose practical value if it no longer matches current selection candidates. Useful warning signs include genetic separation of new families in PCA or genomic relationship analyses, declining accuracy in forward-cycle or leave-family-out validation, introduction of new donor parents or genetic pools, and changes in trait definitions or target environments. Random cross-validation may still look acceptable even when these deployment-focused tests deteriorate.
These signals are more informative than updating the training population on a fixed schedule. A practical review combines genetic connectivity, phenotype comparability, and forward-validation performance. Historical records can remain useful when they continue to connect current candidates, environments, or important alleles.
What should not be over-weighted in a training population?
More records are not automatically better. Poorly phenotyped entries, unresolved sample-identity problems, large groups of nearly redundant siblings, genetically disconnected material, and incompatible historical marker datasets can add little relevant information or create avoidable bias. Before expansion, ask whether each proposed group improves representation of the target population and whether its phenotype and genotype data can be harmonized with the existing training matrix.
Distant or historical material should not be discarded automatically. It may preserve rare alleles, connect long-term trials, or become useful when breeding objectives expand. Its contribution should be evaluated through population diagnostics and validation rather than assumed from sample count alone.
Lifecycle management: updating and pruning the training population
A genomic selection model is not static. Selection changes allele frequencies, recombination reshapes haplotypes, new parents alter population structure, and target environments can shift. Prediction accuracy may therefore decline over time, but the rate of decline is program-specific. Training-population updates should be triggered by evidence of reduced representativeness or forward-prediction performance rather than by a universal schedule.
The rolling training population strategy
A rolling training architecture can help keep the model connected to current candidates while preserving useful historical information:
- Incorporate Advanced Selections: Add newly genotyped and reliably phenotyped lines from relevant trials so the training archive tracks the genetic material currently entering selection decisions.
- Reassess Distant Historical Material: Historical lines should be down-weighted or removed only when they contribute little genetic connectivity or phenotype comparability. Older data can remain valuable when they link environments, families, or important alleles.
- Account for Distinct Genetic Pools: When programs combine genetically distinct populations, model structure should reflect that complexity. In livestock, Misztal et al. (2022) discuss genomic evaluation with multibreed and crossbred data, illustrating why population composition must be represented explicitly rather than treated as a single homogeneous pool.
For protocols on standardizing marker inputs and ensuring multi-year dataset comparability, review the guide on building GS-ready datasets from array and sequencing outputs, or explore marker density guidelines in choosing marker density for breeding cohorts.
Training population readiness checklist
The following checklist is intended as a project review framework rather than a universal set of pass/fail thresholds.
- Genetic Connectivity: Genomic relationship, PCA, pedigree, or haplotype analyses show that the training population adequately represents the candidate population and its major families or genetic pools.
- Phenotyping Rigor: Trait records are generated with a documented trial design, comparable phenotype definitions, appropriate environmental replication where needed, and an estimated level of phenotype reliability suitable for the intended prediction task.
- Genotyping Quality: Marker call rate, missingness, allele-frequency filters, genome-build consistency, and imputation quality are assessed using thresholds validated for the species, platform, and downstream model rather than copied from a different population.
- Training Size: Sample size is evaluated jointly with trait heritability, family structure, effective population size, marker information, and target relatedness. Additional samples should add useful genetic or phenotypic information.
- Cross-Validation Audit: Structure-aware validation, such as leave-family-out or forward-cycle testing, is used when it reflects the intended deployment scenario. Detailed validation designs are covered in the companion guide on cross-validation for genomic prediction in breeding.
Decision-ready deliverables for breeding pipelines
Implementing genomic selection requires translation of genotype and phenotype data into auditable prediction packages that breeding teams can review and reproduce. High-standard deliverables include:
- Curated Training Matrix: Harmonized genotype (G) and phenotype (Y) matrices formatted for standard quantitative genetics engines such as ASReml, rrBLUP, BGLR, or sommer.
- Population Structure Diagnostics: Genomic relationship heatmaps, PCA projections, kinship summaries, and family-level comparisons linking training entries to candidate cohorts.
- Cross-Validation Performance Summary: Empirical prediction accuracy, bias, calibration or slope, and ranking consistency across traits and validation schemes.
- GEBV and Selection Index Tables: Candidate rankings with uncertainty metrics and clearly documented model versions, ready for breeder review rather than presented as guaranteed biological outcomes.
For programs leveraging targeted marker panels or low-pass sequencing to scale candidate genotyping, explore our crop genotyping array services and low-coverage whole-genome sequencing (lc-WGS) solutions, or consult our broader molecular breeding and genotyping workflows.
How CD Genomics can help
As a specialized agricultural genomics and bioinformatics provider, CD Genomics supports plant and animal breeding teams in designing, sequencing, and evaluating genomic selection training populations. Project support can include germplasm sampling design, founder sequencing, high-density array genotyping, low-coverage whole-genome sequencing, genotype harmonization, training-population diagnostics, and customized genomic prediction modeling. The appropriate workflow depends on the breeding population, trait, available historical data, and intended validation scenario. To explore our comprehensive agricultural capabilities, visit the agricultural genomics services overview, or review our integrated genomic breeding services overview. All laboratory and analytical services described herein are provided strictly for Research Use Only (RUO) in agricultural research and breeding programs, and are not intended for clinical or diagnostic use.
Frequently asked questions (FAQ)
Q1: What is the single most important factor determining genomic prediction accuracy?
A: There is no single factor that dominates every breeding dataset, but genetic connectivity between the training population and target candidates is consistently important. Prediction also depends on phenotype reliability, training size, marker information, trait architecture, population structure, and validation design. A genetically related training set with poor phenotypes can still perform poorly.
Q2: How large should a training population be for commercial crop breeding?
A: There is no universal minimum. Closely related families may achieve useful prediction with several hundred well-phenotyped individuals, whereas diverse or weakly connected populations may require substantially larger training sets. The appropriate size should be established empirically by examining learning curves, relatedness, trait reliability, and realistic cross-validation performance.
Q3: Is it better to spend breeding budget on more phenotyping replicates or more genotyped samples?
A: It depends on the current bottleneck. When phenotype reliability is poor, improving trial design, replication, environmental connectivity, or spatial adjustment may provide more value than simply enlarging a noisy training set. When phenotypes are already reliable but candidate diversity is underrepresented, expanding the training population may be more useful. Budget allocation should therefore be evaluated jointly with trait heritability, target environments, training size, and genotyping cost.
Q4: Why does genomic prediction accuracy decline in subsequent breeding generations?
A: Accuracy can decline when new recombination, selection, new parents, or changing environments reduce the similarity between historical training data and current candidates. The relevant warning signal is declining performance in forward or family-aware validation, not the passage of a fixed number of generations. Historical data can remain useful when genetic and phenotypic connectivity is preserved.
Q5: How do algorithmic optimization methods like CDmean improve training set selection?
A: CDmean selects a training subset to maximize the average expected reliability, or coefficient of determination, of predictions for the target candidate set. PEVmean addresses a related but different objective by minimizing average prediction error variance. These methods can reduce redundant sampling and improve efficiency when only part of a large candidate nursery can be phenotyped, but their advantage should be confirmed with population-appropriate validation.
References
- Bazzer, Sumandeep Kaur, Guilherme Oliveira, Jason D. Fiedler, et al. "Genomic Strategies to Facilitate Breeding for Increased β-Glucan Content in Oat (Avena sativa L.)." BMC Genomics, vol. 26, 2025, Article 35.
- Bhattarai, Gehendra, Bo Liu, James Correll, and Ainong Shi. "Genome-Wide Association Study and Genomic Prediction of Leaf Spot (Stemphylium vesicarium) Resistance in Spinach Diversity Panel." Frontiers in Plant Science, vol. 16, 2025, Article 1663650.
- Flores-Saavedra, Martín, Jon Bančič, Yuliza Huamán, Andrea Arrones, Oscar Vicente, Mariola Plazas, Santiago Vilanova, Pietro Gramazio, and Jaime Prohens. "Water Stress Tolerance, Genomic Selection and Identification of Genomic Regions in a MAGIC Population of Eggplant." Theoretical and Applied Genetics, vol. 139, 2026, Article 153.
- Misztal, Ignacy, Yutaka Steyn, and Daniela A. L. Lourenco. "Genomic Evaluation with Multibreed and Crossbred Data." JDS Communications, vol. 3, no. 2, 2022, pp. 156–159.
- Edwards, Stefan McKinnon, Jaap B. Buntjer, Robert Jackson, et al. "The Effects of Training Population Design on Genomic Prediction Accuracy in Wheat." Theoretical and Applied Genetics, vol. 132, no. 7, 2019, pp. 1943–1952.
- Berro, Inés, Bettina Lado, Rafael S. Nalin, Martin Quincke, and Lucía Gutiérrez. "Training Population Optimization for Genomic Selection." The Plant Genome, vol. 12, no. 3, 2019, Article 190028.
Send a MessageFor any general inquiries, please fill out the form below.



