Genomic Selection for Plant and Animal Breeding
Use a genotyped and phenotyped training population to build and validate genomic prediction models, then calculate genomic estimated breeding values (GEBVs) for breeding candidates. CD Genomics connects project assessment, genotyping, data preparation, model validation, candidate ranking, and model-update planning in one solution.
What This Solution Helps You Do
Genomic Selection: What It Does and What You Receive
Many breeding programs already have marker data and phenotype records but still cannot defend a candidate ranking across families, environments, or cycles. The unresolved issue is usually not the absence of one algorithm; it is whether the training evidence represents the candidates and whether validation reflects the way predictions will actually be used.
This solution treats genomic selection as an evidence system: define the target trait and selection unit, audit training and candidate populations, establish a compatible genotype and phenotype foundation, compare prediction models under decision-relevant validation, and translate the results into GEBVs, ranking evidence, and an update plan. For the broader molecular-breeding context, see Molecular Breeding and Genotyping.
How the Results Support Your Breeding Decisions
1. Define the breeding objective
Confirm the target trait, selection unit, breeding population, and decision the model must support.
2. Review project readiness
Assess the training population, future candidates, and available phenotypes, genotypes, pedigree, and environment information.
3. Generate or prepare genotype data
Select SNP arrays, GBS, or low-coverage WGS, or harmonize genotype data that your team already has.
4. Prepare model-ready evidence
Align sample IDs, marker information, phenotypes, relationships, environments, and breeding-cycle metadata.
5. Build and validate prediction models
Compare justified models and test performance, bias, and ranking stability under a validation design that reflects future use.
6. Rank candidates and plan updates
Deliver trait-level GEBVs, candidate-ranking evidence, deployment boundaries, and recommendations for the next model update.
What We Can Do for Your Breeding Program
We can review an existing genomic-selection program, generate a compatible genotyping foundation, build a new prediction model, or update a model for a new cycle. The first review connects your breeding objective, training population, phenotype records, and future candidates so you can see what is usable now and what must be added before ranking.
| What we review | Question it must answer | How it affects the project |
|---|---|---|
| Trait and selection unit | Is the phenotype defined consistently for the line, family, hybrid, individual, or group being ranked? | Confirms what the GEBV represents and prevents incompatible records from being pooled. |
| Training population | Does it capture the genetic backgrounds, trait variation, and relevant environments expected among candidates? | Determines whether model training is credible for the intended deployment population. |
| Candidate population | How closely are future candidates related to the training material, and where could population shift occur? | Defines the validation split and the limits of candidate ranking. |
| Phenotype and environment records | Are units, trial designs, contemporary groups, locations, seasons, and management factors interpretable? | Separates trait signal from avoidable design and metadata effects. |
| Genotype and pedigree resources | Can marker sets, genome builds, sample identities, pedigrees, and batches be aligned? | Shows whether existing data can be reused or a new genotyping layer is required. |
A readiness review can begin with partial data.
Missing fields do not automatically exclude a project. They are recorded as decision-relevant gaps, then addressed through data harmonization, targeted genotyping, revised validation, or a staged pilot. Numerical acceptance criteria are confirmed for the species, population, trait, and data route during project review.
Choose the Genomic Selection Solution That Matches Your Starting Point
You do not need to start with the same data or at the same stage as another breeding program. Choose the route that best describes what you have now. The initial review confirms which inputs can be reused, what must be added, and which result can be supported without overextending the available evidence.
1. You already have genotype and phenotype data
2. You have phenotypes but need new genotyping
3. You already use genomic selection and need a model update
Not sure which route applies?
Begin with the files and sample information you already have. The readiness review separates reusable evidence from gaps and recommends a staged pilot when a full prediction workflow is not yet supportable.
How the Genomic Selection Workflow Works
The workflow moves from a compatible genotype foundation to model-ready data, project-relevant validation, and candidate ranking. Quality checks are placed at each handoff so that sample, marker, phenotype, population, or transfer problems are found before they affect breeding decisions.
Step 1: Select and Prepare the Genotyping Route
The preferred route depends on what is already available, whether a stable marker set exists, how broadly the genome must be represented, and how candidates will be added in later cycles. The routes below are alternatives or complements, not universal service tiers.
| Genotyping method | How it supports genomic selection | When it is a strong fit | What to confirm |
|---|---|---|---|
| SNP arrays or targeted panels | Generate a stable marker backbone for repeat cohorts and model updates. | A relevant panel exists and cross-batch, cross-cycle comparability is a priority. | Fixed content may not capture population-specific or newly discovered variation. |
| Genotyping-by-sequencing (GBS) | Samples reproducible genome fractions to build marker data for large plant or animal cohorts. | The species or population lacks a mature array, or flexible marker discovery is useful. | Missingness, locus consistency, and cross-batch harmonization must be evaluated. |
| Low-coverage WGS with imputation | Produces genome-wide sequence evidence that can be harmonized through a suitable reference and imputation strategy. | Dense coverage and reusable sequence data are important, and reference-panel fit can be assessed. | Prediction quality depends on sequencing QC, reference suitability, and imputation performance. |
How Genotyping Can Be Combined Across Populations and Cycles
The training population and future candidates do not always need identical laboratory workflows, but their marker information must remain compatible for prediction. Any mixed-platform or high-density/low-density design is accepted only after marker overlap, reference suitability, imputation performance, and batch effects have been evaluated.
| Project situation | Practical strategy | What must be demonstrated |
|---|---|---|
| Training data come from more than one platform or batch | Align genome build, marker identity, allele coding, sample identity, and shared marker content before combining records. | Platform or batch differences do not create artificial relationships or unstable candidate rankings. |
| New candidates need routine, repeatable genotyping | Use a stable array, panel, GBS design, or low-coverage WGS workflow that preserves sufficient compatibility with the training marker backbone. | Candidate genotypes can be projected into the same prediction framework used for model training. |
| A high-density reference set supports lower-density routine cohorts | Evaluate a high-density reference plus lower-cost candidate genotyping and imputation design. | The reference represents the target population and imputation accuracy is suitable for the intended ranking decision. |
| New families, cycles, or environments are added later | Track assays, batches, population changes, and new phenotypes, then revalidate before recalibration or retraining. | The updated model remains relevant to the next deployment group rather than only the original training population. |
Explore confirmed same-site routes for Crop Genotyping Array Services, Livestock Genotyping Array Services, and Low-Coverage WGS. The selected route is documented against the target population and update strategy before model development.
Step 2: Prepare Genotype and Phenotype Data
Model comparison is meaningful only after genotype, phenotype, and population evidence has been made internally consistent. The QC plan therefore follows the failure points that could change a candidate ranking.
Phenotype definition and adjustment
Review trait units, repeated records, trial design, contemporary groups, environment structure, and missingness. Adjusted phenotypes or model terms are agreed according to the breeding design rather than imposed as a generic preprocessing step.
Genotype QC and harmonization
Check sample identity, marker orientation, genome build, allele coding, missingness, population-appropriate frequency filters, and batch compatibility. Imputation is evaluated where it changes marker continuity or candidate coverage.
Relationship and population structure
Use genomic relationships, pedigree information when available, and population summaries to identify duplicates, unexpected relatedness, family imbalance, or candidate groups poorly represented by the training set.
Analysis-ready alignment
Create an auditable mapping among sample IDs, phenotypes, genotypes, pedigree, environments, traits, and breeding cycles so every modeled record can be traced back to the source evidence.
Programs with legacy files or multiple data sources can also use Agricultural Genomic Data Analysis to support variant, population, and downstream data preparation.
Step 3: Build and Validate the Prediction Model
GBLUP and Bayesian genomic prediction models provide different assumptions about how genome-wide marker effects are distributed. The useful comparison is not which method is fashionable, but which model remains informative under a validation scenario that resembles the next family, cycle, cohort, or environment.
1. Establish a transparent baseline
Fit a model appropriate to the breeding design, such as GBLUP, and document how genomic relationships, fixed effects, and available pedigree information enter the analysis.
2. Compare justified alternatives
Evaluate Bayesian or other models only when trait architecture, population size, multi-trait information, or environment structure creates a meaningful reason to compare assumptions.
3. Match validation to deployment
Use random, family-aware, generation-aware, cohort-aware, or environment-aware partitions according to the future selection question. A random split alone can be optimistic when close relatives appear in both training and validation sets.
4. Review performance and bias together
Summarize predictive ability or accuracy under an agreed definition, calibration or bias, ranking stability, and uncertainty across folds or scenarios. No single metric is treated as proof that the model will transfer to an unrelated population.
How to Interpret the Results
A GEBV is a statistical prediction for a defined trait, population, model, and evidence base. It supports selection and prioritization; it does not by itself identify a causal variant, prove biological mechanism, or guarantee performance in a new genetic background or environment.
Results and Deliverables for Breeding Decisions
The result package is scoped around what the breeding team needs to review, approve, and update. Exact data structures and compatible formats are confirmed during project scoping rather than assumed in advance.
Project readiness review
A traceable summary of usable inputs, population coverage, missing evidence, and the conditions required to proceed.
Genotype and phenotype QC
Sample-, marker-, trait-, environment-, and relationship-level findings that explain inclusions, exclusions, transformations, and unresolved risks.
Model performance and validation
Side-by-side evidence on prediction performance, bias, stability, and transfer limits under project-relevant validation scenarios.
Trait-level GEBVs and candidate ranking
Predicted breeding values and ranking evidence linked to the target trait, population, model version, and validation context.
Where the model can be used
A clear statement of which candidates, cohorts, cycles, and environments are represented—and where additional evidence is needed.
Model update recommendations
Priorities for adding phenotypes, genotypes, families, cohorts, or environments before the next training and selection cycle.
Scientific Basis
Genome-wide marker prediction was established as a way to estimate total genetic value from dense marker data, and later work formalized genomic relationship approaches and their use in both animal and plant breeding. The literature also shows why training-population design, model assumptions, genotype-by-environment structure, and validation strategy must be reviewed together. Published results demonstrate the framework; they are not performance guarantees for a new project.
Published Research Case: Earlier Quality Selection in Bread Wheat
A commercial bread wheat study shows how genomic data and a faster, lower-cost correlated phenotype can be combined to support earlier selection for quality traits that are expensive to measure directly.
Published Study
Azizinia, S., Mullan, D., Rattey, A., Godoy, J., Robinson, H., Moody, D., Forrest, K., Keeble-Gagnere, G., Hayden, M. J., Tibbits, J. F. G., & Daetwyler, H. D. (2023). Improved multi-trait prediction of wheat end-product quality traits by integrating NIR-predicted phenotypes. Frontiers in Plant Science, 14, 1167221. DOI: 10.3389/fpls.2023.1167221.
Research question
Could NIR-predicted secondary traits improve genomic prediction of six costly end-product quality traits and make selection practical earlier in a commercial wheat breeding cycle?
Study design
The study used 1,400–1,900 bread wheat lines with laboratory quality measurements collected across eight years, plus NIR-predicted records for approximately 27,000 lines. Lines were genotyped with a 40K wheat and barley SNP array and imputed with exome-sequence data. Single- and multi-trait GBLUP models were assessed by fivefold cross-validation and forward prediction across years.
Key findings
Prediction accuracy across the six quality traits ranged from 0.28 to 0.64. Adding NIR-predicted data improved multi-trait prediction overall, but the gain depended on the trait and protein was an exception. Forward prediction also improved as observations from more breeding years were added to the training evidence.
Why it matters for project design
The case connects a costly breeding decision to a practical evidence strategy: use an economical correlated measurement to expand the reference set, verify the genetic relationship with the target trait, and test whether the resulting model improves prediction under a future-cycle validation design.
What This Study Does Not Prove
These results came from specific commercial bread wheat lines, quality assays, NIR predictions, a 40K SNP platform, and the study's cross-validation and forward-prediction designs. The reported accuracies and benefit of a correlated secondary trait should not be assumed to transfer to another population, trait, species, or data-generation route.
Illustration: original conceptual summary based on the cited study. Not a reproduction of the published figure.
Samples and Data Requirements
You can start with existing genotype data or submit genomic DNA for a new genotyping project. A feasibility review can begin before every file is finalized, and route-specific DNA requirements are confirmed before shipment.
| Your starting point | What you can submit | Supporting information |
|---|---|---|
| You already have genotype data | SNP array, GBS, low-coverage WGS, or imputed genotype data with sample and marker identifiers. | Target traits, phenotype records, training/candidate group labels, platform and genome-build information, and pedigree or environment metadata when available. |
| You need new genotyping | High-quality genomic DNA from plant or animal individuals. Required concentration, quantity, purity, and integrity depend on the selected array or sequencing route. | Unique sample IDs, species, variety/line/breed, training or candidate status, cohort/family/batch labels, and the intended breeding decision. |
| You have biological samples but no extracted DNA | Acceptance depends on the plant or animal sample type, preservation method, and extraction route. | Confirm the exact sample type and storage condition during feasibility review before shipping. |
Required Project Information
Additional Information That Helps
Why Choose CD Genomics
FAQ
Start Your Genomic Selection Project
Share your breeding objective, population structure, target traits, and current genotype and phenotype resources. We will use them to identify the next evidence decision—not to prescribe a standard package.
For foundational concepts and model background, read the Genomic Selection in Plant and Animal Breeding Guide.
References
Meuwissen, T. H. E., Hayes, B. J., & Goddard, M. E. (2001). Prediction of total genetic value using genome-wide dense marker maps. Genetics, 157(4), 1819–1829. DOI: 10.1093/genetics/157.4.1819.
VanRaden, P. M. (2008). Efficient methods to compute genomic predictions. Journal of Dairy Science, 91(11), 4414–4423. DOI: 10.3168/jds.2007-0980.
Crossa, J., Pérez-Rodríguez, P., Cuevas, J., Montesinos-López, O., Jarquín, D., de los Campos, G., et al. (2017). Genomic Selection in Plant Breeding: Methods, Models, and Perspectives. Trends in Plant Science, 22(11), 961–975. DOI: 10.1016/j.tplants.2017.08.011.
Hickey, J. M., Chiurugwi, T., Mackay, I., Powell, W., & Implementing Genomic Selection in CGIAR Breeding Programs Workshop Participants. (2017). Genomic prediction unifies animal and plant breeding programs to form platforms for biological discovery. Nature Genetics, 49, 1297–1303. DOI: 10.1038/ng.3920.
Azizinia, S., Mullan, D., Rattey, A., Godoy, J., Robinson, H., Moody, D., Forrest, K., Keeble-Gagnere, G., Hayden, M. J., Tibbits, J. F. G., & Daetwyler, H. D. (2023). Improved multi-trait prediction of wheat end-product quality traits by integrating NIR-predicted phenotypes. Frontiers in Plant Science, 14, 1167221. DOI: 10.3389/fpls.2023.1167221.
All products and services are For Research Use Only and not for diagnostic or therapeutic use.
For research purposes only, not intended for clinical diagnosis, treatment, or individual health assessments.
Send a MessageFor any general inquiries, please fill out the form below.