Agricultural genomics resource banner
Cross-Validation for Genomic Prediction: Will Your Model Work in the Next Breeding Cycle?

Cross-Validation for Genomic Prediction: Will Your Model Work in the Next Breeding Cycle?

Genomic prediction cross-validation strategies

In quantitative genetics and commercial breeding pipelines, statistical models are routinely evaluated by cross-validation predictive ability or prediction accuracy. Yet a model that performs strongly under random cross-validation can show substantially lower predictive ability when deployed on new families, future breeding cycles, or previously unseen environments. This discrepancy often reflects a mismatch between the validation design and the intended deployment scenario, although model specification, phenotype quality, population structure, and genotype-by-environment (G × E) effects can also contribute. When validation folds contain relationships, environments, or preprocessing information that would not be available at deployment, the reported performance can be optimistic. This guide compares practical cross-validation strategies, explains family and environmental leakage, distinguishes predictive ability from calibration and ranking performance, and provides a deployment-oriented framework for evaluating genomic prediction models before operational breeding decisions.

Key takeaways

  • Random k-fold cross-validation can be optimistic when close relatives or repeated genetic material are distributed across training and validation folds but the intended deployment involves new families or more distant candidates.
  • The validation scheme should mirror the breeding decision: CV1 for untested genotypes in observed environments, CV2 for partially tested genotypes, CV0 for observed genetic material in new environments, CV00 for new genotypes in new environments, or forward validation for future breeding cycles.
  • High Pearson correlation (r) does not guarantee unbiased selection; models must be evaluated for regression slope (β ≈ 1.0), prediction error variance, and top-decile candidate ranking consistency.
  • Forward-in-time validation, in which earlier cycles or years predict a later cycle, often provides a more deployment-relevant estimate of future-cycle predictive performance.
  • Family-aware or cluster-aware validation is useful when the operational question is whether a model can transfer beyond the close relationships represented in the training population.

Why random cross-validation can overestimate breeding accuracy

Which cross-validation strategy matches your breeding objective? The answer depends on the operational stage at which selection candidates will be evaluated. If the goal is to predict progeny from new biparental crosses, withholding entire families or crosses provides a more realistic test than randomly masking individuals. If the intended use is prediction within the same connected breeding population, however, a random split may still be informative. Validation should reproduce the genetic distance, environmental novelty, and data availability expected at deployment.

For strategic principles on selecting individuals and sizing training cohorts before cross-validation, explore the dedicated guide on training population design for genomic selection.

Cross-validation schemes for breeding programs

The family leakage trap

In standard random 5-fold or 10-fold cross-validation, genotyped individuals are assigned to folds without regard to pedigree or genetic clusters. In breeding populations containing large full-sib or half-sib families, this can place close relatives in both training and validation sets. Part of the apparent predictive ability may then arise from close genomic relationships between the two sets. That signal can be useful for within-family prediction, but it may not transfer to newly introduced crosses or genetically distant candidates. When the deployment target is a new family, leave-family-out or cluster-aware validation provides a more stringent estimate of transferability.

Environmental and year confounding

In multi-environment trials (MET), a validation design can also become optimistic when the same genotype, closely related material, or strongly correlated environment information appears on both sides of the split in a way that would not occur at deployment. The model may then benefit from environmental or genetic information that is unavailable when predicting a future site, year, or family. Environment-aware partitioning is therefore important when the breeding objective includes extrapolation across years, locations, or management conditions.

When random cross-validation is appropriate

Random cross-validation is not inherently incorrect. It can be appropriate when the intended deployment closely resembles the current dataset—for example, ranking additional candidates from the same connected breeding population, comparing algorithms under a common data distribution, or filling missing phenotypes among material that shares similar relatedness and environment structure. Problems arise when random-fold performance is interpreted as evidence that the model will transfer to new families, future generations, or unobserved environments. In those cases, the validation split should deliberately reproduce the expected deployment gap. A useful rule is to ask which sources of information will still be available when the model is used operationally, and then prevent the validation design from using anything beyond that information.

Cross-validation paradigms in breeding

To obtain realistic estimates of model performance, quantitative geneticists categorize cross-validation into four structured schemes based on genotype status, family structure, and environmental exposure.

For an overview of statistical foundations and quantitative modeling workflows, review the guide to genomic selection in plant and animal breeding.

Cross-validation scheme Partitioning strategy Simulated breeding decision Typical interpretation Primary risk / Limitation
Random k-Fold CV Random assignment across all individuals Prediction within a genetically connected population Useful for matched-distribution or within-population prediction, but potentially optimistic for new-family deployment May overstate transferability when close relatives span folds
Leave-Family-Out (LFO) / Cluster-Aware CV Entire biparental crosses or PCA clusters withheld Predicting progeny from unphenotyped parental crosses More stringent when the intended task is prediction of new crosses or genetic clusters Requires multi-family cohorts; sensitive to family-specific QTLs
Multi-Environment CV (CV1 / CV2 / CV0) Stratified by location, year, or management tier Predicting lines across untested locations or future years Strongly dependent on G × E structure and environmental relatedness G × E interactions can degrade across-location predictions
Forward-in-Time (Cycle-to-Cycle) CV Train on generations 1…t; validate on generation t+1 Commercial deployment to select the next breeding cycle More representative of future-cycle deployment when historical data are available Requires multi-year historical trial data and consistent genotyping

Multi-environment cross-validation schemes (CV1, CV2, CV0, CV00)

Multi-environment genomic prediction studies use several validation labels. Definitions are not completely uniform across the literature, so the deployment scenario should always be described explicitly rather than relying on a label alone. A commonly used interpretation is:

  • CV1 (Untested Genotypes in Tested Environments): Target lines have not been phenotyped, but the target environments are represented in the training data. This evaluates prediction of new genetic material under environments already observed by the model.
  • CV2 (Partially Tested Genotypes in Tested Environments): Target lines have records in some environments and are predicted in others. This reflects sparse multi-environment testing where information is borrowed across locations or years.
  • CV0 (Tested Genotypes in Untested Environments): Genetic material represented in the training archive is predicted in a new environment, year, or location that was not observed during model fitting.
  • CV00 (Untested Genotypes in Untested Environments): Both the target genotypes and target environments are new to the model. This is a stringent extrapolation scenario and should be distinguished from CV0 when that distinction is used in the source study.

Burgueño et al. (2012) formalized CV1 and CV2 for multi-environment genomic prediction in wheat, while Khanna et al. (2022) used CV1, CV2, and CV0 to evaluate 17 years of historical rice drought-breeding data and showed that predictive ability depended strongly on whether target environments had been observed previously. These studies illustrate why validation design should be selected around the exact breeding question rather than around a single headline accuracy value.

Choose validation by deployment question

  • New lines in known environments: Use CV1 or a closely matched genotype-holdout design.
  • Partially tested lines across the current trial network: Use CV2.
  • Known material in a new year or location: Use CV0 or an explicit environment-holdout design.
  • New genetic material in a new environment: Use CV00 when that terminology is appropriate, and report the exact holdout design.
  • New family or cross: Use leave-family-out or cluster-aware validation.
  • Next breeding cycle: Use forward-in-time validation that trains only on information available before the target cycle.

Forward-in-time validation for future breeding cycles

Forward validation orders training and testing data chronologically instead of randomly. This is particularly useful when a breeding program wants to know how a model trained on historical cycles may perform on the next cycle. In oat, Bazzer et al. (2025) used historical breeding data to predict a later year and found that multi-year training information produced more useful forward predictions than relying on a single historical year. The exact predictive ability remains trait- and population-dependent, but the design is valuable because it reproduces the temporal information boundary faced in operational breeding. Similarly, Bhattarai et al. (2025) showed that prediction performance in spinach depended on the validation population, reinforcing that strong within-panel results do not automatically imply equally strong transfer to different genetic material.

Genomic prediction calibration and inflation

Beyond Pearson correlation: evaluating bias, inflation, and ranking

Reporting only the Pearson correlation coefficient (ry,ŷ) between predicted breeding values (GEBV) and observed phenotypes is insufficient for practical breeding selection. A model can have a high correlation but suffer from severe scale inflation or rank distortion in the top selection bracket.

Prediction accuracy versus predictive ability

In quantitative genetics, predictive ability is often reported as the correlation between predicted genetic values (ĝ) and observed or adjusted phenotypic values (y): ry,ĝ. Under common quantitative-genetic assumptions, predictive ability can be converted to an approximate prediction accuracy by accounting for phenotype reliability or an appropriate heritability term:

rg,ĝ = ry,ĝ / √(h2)

This normalization is an approximation rather than a universal conversion rule. The reliability or heritability term must match the phenotype definition, genetic target, and trial model used in the analysis; realized accuracy can also be affected by population structure, training-target relatedness, G × E, and model specification.

Assessing model bias and inflation (β)

Model calibration is evaluated by regressing observed phenotypes on predicted breeding values (y = α + β × ĝ). The regression slope (β) indicates whether the model's predictions are appropriately scaled:

  • β ≈ 1.0: Predictions are approximately well calibrated in scale, subject to sampling uncertainty and the reliability of the validation target.
  • β < 1.0: Predictions may be inflated in scale, with predicted differences among candidates larger than supported by the validation data.
  • β > 1.0: Predictions may be deflated in scale, with candidate differences compressed relative to the validation target.

Top-decile ranking and coincidence index

Breeders do not advance entire populations; they select the top 5%–10% of candidates. Therefore, model validation must evaluate ranking fidelity in the upper tail:

  • Spearman Rank Correlation: Evaluates monotonic candidate order, which is robust against non-linear scaling distortions.
  • Coincidence Index (CI): Measures the proportion of truly top-performing individuals recovered within a predicted selection bracket. The result should be interpreted relative to random selection and to the recovery rate required by the breeding program rather than against a universal cutoff.

Preventing data leakage during dataset preparation

Data leakage occurs when information from the validation set inadvertently influences the training phase, artificially boosting validation scores. Common sources of leakage include:

  • Genotype Preprocessing Leakage: Phasing, imputation, dimensional reduction, or other genotype preprocessing can become optimistic if validation individuals contribute information that would not be available when the model is deployed. The preprocessing workflow used during validation should mimic the intended production pipeline. For reference-panel design principles, consult how to build a genotype imputation reference panel for crop breeding.
  • GWAS Feature Selection Leakage: Performing GWAS marker discovery on the entire dataset and selecting top-associated SNPs before cross-validation. Marker discovery must be executed strictly within each training fold.
  • Phenotype-Adjustment Leakage: Spatial, block, environmental, or feature-adjustment models should be fitted so that validation outcomes do not influence parameters used to evaluate the genomic prediction model. The exact implementation should mirror how adjusted phenotypes will be generated in the operational pipeline.

Standardized workflows for formatting marker matrices and eliminating leakage risks prior to quantitative modeling are detailed in building GS-ready datasets from array and sequencing outputs. Guidelines for selecting cost-effective genotyping densities are covered in choosing marker density for breeding cohorts.

Genomic prediction validation checklist

Genomic prediction model validation checklist

  • Deployment Alignment: The validation scheme explicitly matches the future decision: within-population, new-family, new-environment, new-cycle, or combined extrapolation.
  • Family & Clone Structure: Family isolation matches the intended deployment. Entire families should be withheld when testing transfer to new crosses, while within-family validation may retain related material when that reflects operational use.
  • Feature Discovery Independence: GWAS pre-selection, supervised feature selection, and learned preprocessing steps are conducted within training data when the same information would be unavailable for future candidates.
  • Accuracy Definition: Predictive ability, approximate prediction accuracy, phenotype reliability, and any heritability adjustment are reported with clear definitions rather than treated as interchangeable metrics.
  • Calibration Slope Audit: Calibration slopes are interpreted relative to 1.0 together with uncertainty, validation-set size, and reliability of the observed target values rather than against a fixed universal pass/fail range.
  • Tail-Selection Metrics: Coincidence Index, Spearman rank correlation, and recovery of the top selection fraction are interpreted against random selection and the breeder's operational selection intensity.
  • Crossbred & Multibreed Compatibility: When animal or multibreed breeding data are analyzed, validation should reflect breed composition and the relevant reference population. Misztal et al. (2022) emphasized that crossbred prediction accuracy depends strongly on whether the reference population represents the target breed type.

To explore high-throughput genotyping options for large-scale candidate screening, review our crop genotyping array services and low-coverage whole-genome sequencing (lc-WGS) solutions, or explore our end-to-end molecular breeding and genotyping workflows.

How CD Genomics can help

Validating genomic prediction models requires robust statistical frameworks, multi-environment data integration, and validation folds that reflect the intended deployment. CD Genomics provides specialized quantitative genetics and bioinformatics consulting to help agricultural breeding programs design realistic cross-validation schemes, detect hidden data leakage, calibrate genomic prediction algorithms, and evaluate multi-generation forward-cycle predictive performance. We deliver fully audited analytical packages and decision-ready GEBV ranking tables tailored to your program's commercial crossing cycles. To review our end-to-end analytical capabilities, visit the agricultural genomic data analysis page, or explore our full suite of solutions via the agricultural genomics services overview and genomic breeding services overview. All analytical and sequencing workflows described in this guide are provided strictly for Research Use Only (RUO) in plant and animal agricultural genetics, and are not intended for clinical or diagnostic use.

Frequently asked questions (FAQ)

Q1: Is random 10-fold cross-validation unreliable for commercial crop breeding?
A: Not inherently. Random CV can be informative when future candidates come from the same connected population represented in the current dataset. It becomes optimistic when close relatives are distributed across folds but the intended deployment involves new families, future cycles, or more distant genetic material. In those cases, family-aware or forward validation better matches the operational question.

Q2: What is the difference between CV1, CV0, and CV00 cross-validation?
A: In a commonly used multi-environment framework, CV1 predicts untested genotypes in environments already represented in training. CV0 predicts previously represented genetic material in a new environment or year. CV00 is used in some studies for the more stringent case in which both the genotype and the environment are new. Because terminology can vary among studies, the exact holdout design should always be reported.

Q3: How do I know if my genomic prediction model is suffering from prediction inflation?
A: A common diagnostic is to regress an appropriate observed or adjusted validation target on predicted breeding values. A calibration slope below 1.0 can indicate inflation in the scale of predictions, whereas a slope above 1.0 can indicate deflation. The slope should be interpreted with its uncertainty, validation-set size, and phenotype reliability rather than against a single universal cutoff.

Q4: How should cross-validation be structured when training on multi-year historical data?
A: When the intended use is prediction of a future breeding cycle, forward-in-time validation is usually the most deployment-relevant design. Train using only data available before the target year or cycle, then evaluate candidates in the subsequent period. If environments also change, report whether the validation represents a new year, new location, new genotype set, or a combination of these factors.

Q5: Can GWAS feature selection be performed before setting up cross-validation folds?
A: No. Pre-selecting markers using the entire dataset causes data leakage, artificially inflating cross-validation accuracy. Feature selection and GWAS must be executed strictly within the training fold of each cross-validation iteration.

References

  1. Bazzer, Sumandeep Kaur, Guilherme Oliveira, Jason D. Fiedler, et al. "Genomic Strategies to Facilitate Breeding for Increased β-Glucan Content in Oat (Avena sativa L.)." BMC Genomics, vol. 26, 2025, Article 35.
  2. Bhattarai, Gehendra, Bo Liu, James Correll, and Ainong Shi. "Genome-Wide Association Study and Genomic Prediction of Leaf Spot (Stemphylium vesicarium) Resistance in Spinach Diversity Panel." Frontiers in Plant Science, vol. 16, 2025, Article 1663650.
  3. Misztal, Ignacy, Y. Steyn, and Daniela A. L. Lourenco. "Genomic Evaluation with Multibreed and Crossbred Data." JDS Communications, vol. 3, no. 2, 2022, pp. 156–159.
  4. Burgueño, Juan, Gustavo de los Campos, Kent Weigel, and José Crossa. "Genomic Prediction of Breeding Values when Modeling Genotype × Environment Interaction Using Pedigree and Dense Molecular Markers." Crop Science, vol. 52, no. 2, 2012, pp. 707–719.
  5. Khanna, Apurva, Mahender Anumalla, Margaret Catolos, Sankalp Bhosale, Diego Jarquin, and Waseem Hussain. "Optimizing Predictions in IRRI's Rice Drought Breeding Program by Leveraging 17 Years of Historical Data and Pedigree Information." Frontiers in Plant Science, vol. 13, 2022, Article 983818.
For research purposes only, not intended for clinical diagnosis, treatment, or individual health assessments.
Send a MessageSend a Message

For any general inquiries, please fill out the form below.

For research purposes only, not intended for clinical diagnosis, treatment, or individual health assessments.
We provide the best service according to your needs Contact Us