Polyploid Genotyping and Allele Dosage Analysis
A diploid genotype call cannot distinguish the full set of allele-copy states in tetraploid, hexaploid, or mixed-ploidy breeding material. We design array- or sequencing-based polyploid genotyping workflows, estimate allele dosage with uncertainty, challenge calls with population and technical controls, and deliver dosage-aware datasets for genetic linkage maps, GWAS, and breeding analysis.
What This Solution Helps You Decide
Polyploid Genotyping and Allele Dosage: What It Does
At a biallelic locus, a diploid sample has three genotype states. A tetraploid can have five allele-dosage states, while a hexaploid can have seven. Collapsing every heterozygote into one class removes information that may affect linkage, relatedness, association, and breeding-value models. The first project decision is therefore not which caller to run; it is whether the biological and technical evidence supports dosage-aware analysis.
What usually goes wrong?
What this Solution produces
This page focuses on population genotyping and allele dosage. If the primary question is genome sequencing, assembly, subgenome reconstruction, or multi-omics, begin with Polyploid Genomes in Plants.
What We Can Analyze in Polyploid Breeding Material
A technically clean signal can still be biologically misclassified when ploidy, inheritance, or subgenome behavior is assumed incorrectly. We review the evidence available for the species, parents, population, and reference before fixing the genotype state space.
Ploidy evidence
Known chromosome counts, flow-cytometry results, published cytogenetics, expected cross design, and genome-wide allele-ratio patterns are reviewed together. Mixed or uncertain groups are flagged rather than forced into one model.
Auto- or allopolyploid context
The expected pairing behavior and degree of subgenome divergence affect mapping, marker interpretation, segregation expectations, and whether homeolog-specific filtering is possible.
Reference suitability
Assembly level, haplotype or subgenome resolution, annotation version, repeat content, and alignment ambiguity determine which loci can support dosage calls.
Population design
Parents, progeny, pedigrees, diversity panels, clones, and technical replicates provide different priors and validation opportunities. The analytical design must reflect the actual population.
When this Solution is not the right starting point
Pause dosage calling when the sample identity is unresolved, the biological material contains uncontrolled tissue mixtures, ploidy evidence conflicts across the cohort, the reference cannot distinguish the required subgenomes, or the planned downstream method accepts only presence/absence markers and gains no value from dosage. The feasibility review defines whether added controls, a revised platform, or genome-level work is required first.
Choose the Polyploid Genotyping Solution You Need
The right solution combines the required result—hard dosage calls, genotype probabilities, continuous allele ratios, or a hybrid output—with a laboratory route that can resolve those states in your species and population.
Choose the Dosage Result Needed for the Next Analysis
Not every project should end with a single hard call per locus. The useful representation depends on signal separation, coverage, population design, and what the linkage, association, or prediction model can consume.
Discrete allele dosage
Use when: dosage classes are well separated and supported by platform signal, depth, population priors, or controls.
Decision value: retains nulliplex-to-full-dosage states for interpretable inheritance and dosage-aware models.
Boundary: uncertain calls remain missing or flagged; they are not rounded into a confident class.
Genotype probabilities
Use when: several dosage states remain plausible, especially with low or uneven sequencing depth.
Decision value: preserves uncertainty for downstream tools that can use posterior probabilities or dosage expectations.
Boundary: probability-based results require compatible downstream software and documented priors.
Continuous allele ratio
Use when: fixed dosage categories are poorly supported or a downstream method can model continuous allele fractions.
Decision value: avoids overconfident classification while retaining more signal than a binary presence/absence call.
Boundary: ratios remain sensitive to mapping, amplification, probe, and batch bias.
Why pseudo-diploid calls are a controlled fallback
Combining all heterozygous dosage states into one class can simplify software compatibility, but it discards copy-number information. We use that representation only when it matches the downstream question or as an explicit comparison—not as the silent default for a polyploid cohort.
Match the Laboratory Route to the Dosage Question
Platform selection is based on whether markers already exist, how many individuals must be compared, whether subgenome-specific alignment is possible, and how much signal is needed to distinguish neighboring dosage states.
Crop SNP arrays
Best fit: established markers, repeat cohorts, and species where probe behavior and dosage clustering can be calibrated.
Key evidence: raw allele-intensity files, cluster separation, reference samples, and marker annotation—not genotype labels alone.
GBS or reduced-representation sequencing
Best fit: flexible marker discovery across many samples when a fixed array is unavailable or unsuitable.
Key evidence: allele depths, locus consistency, missingness, mapping uniqueness, and population-aware genotype probabilities.
Whole-genome sequencing
Best fit: broad variant discovery, reference improvement, or projects requiring loci outside existing panels.
Key evidence: depth distribution, allele balance, mapping quality, subgenome ambiguity, and validated variant scope.
Targeted confirmation
Best fit: smaller candidate-marker sets, orthogonal checks, or conversion of validated loci into an operational breeding assay.
Key evidence: assay specificity, dosage-cluster behavior, positive controls, and concordance with discovery data.
A route can be staged
A pilot can compare marker recovery and dosage separation before the full cohort is committed. An array can anchor repeat cohorts, GBS can extend marker discovery, WGS can resolve difficult regions or supply reference evidence, and targeted assays can confirm selected loci. The combination is justified only when genome coordinates, alleles, and sample identities can be reconciled.
Available routes include Crop Genotyping Array Services, Genotyping by Sequencing, and Whole Genome Sequencing. Existing crop-specific options include Oat Genotyping Arrays and Wheat Liquid Phase Genotyping Arrays.
How Polyploid Dosage Calling and Quality Control Work
Raw array intensities or sequence allele counts are converted into dosage evidence only after ploidy, locus behavior, sample identity, uncertainty, segregation, and population consistency have been reviewed.
Step 1: Convert Raw Signals into Allele-Dosage Evidence
The workflow is organized around decisions that can stop or redirect the project. Samples and markers do not advance simply because a software command completed.
1. Freeze the biological assumptions
Record expected ploidy, auto- or allopolyploid context, population design, parents, reference assembly, target loci, and downstream data requirements.
2. Review sample and platform evidence
Audit identity, DNA or source-data QC, replicate design, array intensity or allele depth, coverage distribution, missingness, and batch structure.
3. Build the analyzable marker set
Align loci to the agreed reference, remove or flag ambiguous mappings, reconcile alleles, assess homeolog specificity, and retain the signal fields required for dosage estimation.
4. Estimate dosage and uncertainty
Fit the agreed signal or read-count model using suitable ploidy, error, bias, dispersion, frequency, family, or population assumptions; retain posterior evidence where appropriate.
5. Test biological consistency
Evaluate parents and progeny, expected segregation, replicate concordance, ploidy controls, allele-frequency behavior, and inconsistencies that may indicate sample or locus problems.
6. Release for the stated use
Apply project-specific thresholds, compare alternative representations where needed, document exclusions, and create a dosage-aware handoff for the agreed downstream model.
Step 2: Apply Polyploid-Specific Quality Control
A high call rate can hide systematic dosage errors. Quality review must show whether adjacent dosage states separate, whether the expected population supports those states, and whether technical artifacts concentrate in particular samples, loci, subgenomes, or batches.
Five evidence gates before release
What happens when the evidence is uneven?
We may retain genotype probabilities instead of hard calls, increase the missing state for ambiguous samples or loci, analyze ploidy groups separately, add parents or reference controls, revise mapping filters, compare a continuous representation, increase data support for a subset, or exclude unsupported regions. The outcome is a documented decision, not a silent conversion of uncertainty into certainty.
Deliverables That Preserve Dosage, Uncertainty, and Provenance
The delivery package is designed so a downstream analyst can identify what was observed, what was inferred, which model and version were used, and where the evidence stops.
Dosage-aware VCF
Ploidy-aware genotype fields with agreed allele-depth, depth, quality, dosage, probability, and filter annotations retained where supported.
Analysis matrices
Discrete dosage, expected dosage, genotype-probability, continuous-ratio, or controlled pseudo-diploid matrices prepared for compatible downstream tools.
Sample and marker QC
Missingness, depth or intensity behavior, confidence, heterozygosity, dosage-state distribution, duplicates, batches, and exclusion decisions.
Population consistency evidence
Parent-offspring checks, segregation summaries, allele-frequency review, ploidy-group behavior, and unresolved biological conflicts.
Marker annotation and provenance
Genome build, coordinates, alleles, subgenome or homeolog status where supported, platform source, caller/model, parameters, and release version.
Downstream readiness record
A clear recommendation for which dataset can enter linkage mapping, association, genomic prediction, diversity analysis, or targeted marker validation.
Validated dosage datasets can support Genetic Linkage Map construction, GWAS Services, or broader Agricultural Genomic Data Analysis. Each downstream use receives the representation and filter boundary it can actually support.
Published Research Case: Dosage-Aware Trait Mapping in Hexaploid Sweet Potato
Haque, E., Shirasawa, K., Suematsu, K., Tabuchi, H., Isobe, S., & Tanaka, M. (2023). Polyploid GWAS reveals the basis of molecular marker development for complex breeding traits including starch content in the storage roots of sweet potato. Frontiers in Plant Science, 14, 1181909. DOI: 10.3389/fpls.2023.1181909.
Research question
Could genome-wide SNP evidence be used in an autohexaploid sweet-potato population to investigate starch content alongside dry matter, storage-root weight, and anthocyanin-related variation?
Study design
The researchers evaluated a 204-member F1 population from two contrasting parents, collected field phenotypes across 2019 and 2020, generated ddRAD sequencing data, retained 90,222 SNPs after alignment and filtering, estimated allele dosage under a six-copy model, and applied polyploid GWAS.
Key findings
The study identified trait-associated signals across the full and phenotype-defined subsets, highlighted a starch-content signal consistent across years, and examined candidate genes in the associated homologous group. The important project lesson is the chain from population design and allele-depth evidence to dosage estimation, trait association, cross-year consistency, and candidate follow-up.
Why it matters for this Solution
What this study does not prove
Its marker effects, thresholds, loci, and performance are specific to the studied cross, reference, sequencing data, traits, environments, and analysis. They do not establish expected call quality or marker transferability for another species, population, platform, or breeding program.
Illustration: original conceptual summary based on the cited study. Not a reproduction of the published figure.
Samples, Data, and Project Information We Can Review
The project can start from existing genotype evidence or from biological material requiring a new data-generation route. Exact sample specifications are confirmed only after the species, tissue, platform, ploidy, and downstream objective are reviewed.
Existing array or sequencing data
Extracted DNA or biological samples
Biological and downstream context
Information that changes the project route
Tell us whether ploidy varies among samples, whether parent and progeny data are available, whether raw intensity or allele-depth fields have been retained, whether the reference distinguishes subgenomes, and which downstream software or file standard must be supported. These details determine whether a pilot, new genotyping, data rescue, or full-cohort dosage analysis is appropriate.
Why CD Genomics
FAQ
Discuss Your Polyploid Genotyping and Dosage Project
Start with the crop, expected ploidy, population design, current data or sample type, and the decision the final genotype matrix must support. We will map those inputs to a pilot, data-rescue, or full-cohort workflow.
For platform-selection context, see Genotyping Arrays for Polyploid Crops in Wheat, Oat, Brassica, and Cotton.
References
Voorrips, R. E., Gort, G., & Vosman, B. (2011). Genotype calling in tetraploid species from bi-allelic marker data using mixture models. BMC Bioinformatics, 12, 172. DOI: 10.1186/1471-2105-12-172.
Gerard, D., Ferrão, L. F. V., Garcia, A. A. F., & Stephens, M. (2018). Genotyping Polyploids from Messy Sequencing Data. Genetics, 210(3), 789–807. DOI: 10.1534/genetics.118.301468.
Clark, L. V., Lipka, A. E., & Sacks, E. J. (2019). polyRAD: Genotype Calling with Uncertainty from Sequencing Data in Polyploids and Diploids. G3: Genes, Genomes, Genetics, 9(3), 663–673. DOI: 10.1534/g3.118.200913.
Yamamoto, E., Shirasawa, K., Kimura, T., Monden, Y., Tanaka, M., & Isobe, S. (2020). Genetic Mapping in Autohexaploid Sweet Potato with Low-Coverage NGS-Based Genotyping Data. G3: Genes, Genomes, Genetics, 10(8), 2661–2670. DOI: 10.1534/g3.120.401433.
Haque, E., Shirasawa, K., Suematsu, K., Tabuchi, H., Isobe, S., & Tanaka, M. (2023). Polyploid GWAS reveals the basis of molecular marker development for complex breeding traits including starch content in the storage roots of sweet potato. Frontiers in Plant Science, 14, 1181909. DOI: 10.3389/fpls.2023.1181909.
All products and services are For Research Use Only and not for diagnostic or therapeutic use.
For research purposes only, not intended for clinical diagnosis, treatment, or individual health assessments.
Send a MessageFor any general inquiries, please fill out the form below.