Polyploid Genotyping and Allele Dosage Analysis

A diploid genotype call cannot distinguish the full set of allele-copy states in tetraploid, hexaploid, or mixed-ploidy breeding material. We design array- or sequencing-based polyploid genotyping workflows, estimate allele dosage with uncertainty, challenge calls with population and technical controls, and deliver dosage-aware datasets for genetic linkage maps, GWAS, and breeding analysis.

What This Solution Helps You Decide

Confirm the ploidy and inheritance model Choose arrays, GBS, WGS, or a targeted route Retain dosage probability when hard calls are uncertain Release only markers supported for downstream analysis

Polyploid samples, allele-copy states, dosage-aware genotyping, and downstream validation

Polyploid Genotyping and Allele Dosage: What It Does

At a biallelic locus, a diploid sample has three genotype states. A tetraploid can have five allele-dosage states, while a hexaploid can have seven. Collapsing every heterozygote into one class removes information that may affect linkage, relatedness, association, and breeding-value models. The first project decision is therefore not which caller to run; it is whether the biological and technical evidence supports dosage-aware analysis.

What usually goes wrong?

  • Unverified ploidy: the same copy-number model is applied to accessions, progeny, or loci that may not share one ploidy level.
  • Collapsed heterozygotes: simplex, duplex, triplex, and higher-copy states are reduced to one diploid-like class.
  • Ambiguous signal: limited depth, allele bias, probe behavior, overdispersion, or outliers blur neighboring dosage clusters.
  • Homeolog confusion: reads or probes from related subgenomes create false variants or distorted allele ratios.
  • Unsupported handoff: dosage calls are delivered without probabilities, rejection rules, segregation checks, or downstream-use boundaries.

What this Solution produces

  • A documented ploidy and inheritance assumption for each project group.
  • A platform route matched to the population, marker resources, and required dosage resolution.
  • Discrete dosages, genotype probabilities, continuous ratios, or a justified combination.
  • Polyploid-specific sample, marker, dosage-cluster, and population QC evidence.
  • A versioned VCF and analysis matrix with clear release, exclusion, and uncertainty rules.

This page focuses on population genotyping and allele dosage. If the primary question is genome sequencing, assembly, subgenome reconstruction, or multi-omics, begin with Polyploid Genomes in Plants.

What We Can Analyze in Polyploid Breeding Material

A technically clean signal can still be biologically misclassified when ploidy, inheritance, or subgenome behavior is assumed incorrectly. We review the evidence available for the species, parents, population, and reference before fixing the genotype state space.

Ploidy evidence

Known chromosome counts, flow-cytometry results, published cytogenetics, expected cross design, and genome-wide allele-ratio patterns are reviewed together. Mixed or uncertain groups are flagged rather than forced into one model.

Auto- or allopolyploid context

The expected pairing behavior and degree of subgenome divergence affect mapping, marker interpretation, segregation expectations, and whether homeolog-specific filtering is possible.

Reference suitability

Assembly level, haplotype or subgenome resolution, annotation version, repeat content, and alignment ambiguity determine which loci can support dosage calls.

Population design

Parents, progeny, pedigrees, diversity panels, clones, and technical replicates provide different priors and validation opportunities. The analytical design must reflect the actual population.

When this Solution is not the right starting point

Pause dosage calling when the sample identity is unresolved, the biological material contains uncontrolled tissue mixtures, ploidy evidence conflicts across the cohort, the reference cannot distinguish the required subgenomes, or the planned downstream method accepts only presence/absence markers and gains no value from dosage. The feasibility review defines whether added controls, a revised platform, or genome-level work is required first.

Choose the Polyploid Genotyping Solution You Need

The right solution combines the required result—hard dosage calls, genotype probabilities, continuous allele ratios, or a hybrid output—with a laboratory route that can resolve those states in your species and population.

Choose the Dosage Result Needed for the Next Analysis

Not every project should end with a single hard call per locus. The useful representation depends on signal separation, coverage, population design, and what the linkage, association, or prediction model can consume.

Discrete allele dosage

Use when: dosage classes are well separated and supported by platform signal, depth, population priors, or controls.

Decision value: retains nulliplex-to-full-dosage states for interpretable inheritance and dosage-aware models.

Boundary: uncertain calls remain missing or flagged; they are not rounded into a confident class.

Genotype probabilities

Use when: several dosage states remain plausible, especially with low or uneven sequencing depth.

Decision value: preserves uncertainty for downstream tools that can use posterior probabilities or dosage expectations.

Boundary: probability-based results require compatible downstream software and documented priors.

Continuous allele ratio

Use when: fixed dosage categories are poorly supported or a downstream method can model continuous allele fractions.

Decision value: avoids overconfident classification while retaining more signal than a binary presence/absence call.

Boundary: ratios remain sensitive to mapping, amplification, probe, and batch bias.

Polyploid allele-copy states compared with dosage probabilities, continuous ratios, and collapsed diploid calls

Why pseudo-diploid calls are a controlled fallback

Combining all heterozygous dosage states into one class can simplify software compatibility, but it discards copy-number information. We use that representation only when it matches the downstream question or as an explicit comparison—not as the silent default for a polyploid cohort.

  • Compare information retained by each representation.
  • Define how uncertain and missing calls are encoded.
  • Keep the original signal or read-depth evidence traceable.
  • Record which downstream analyses accept each dataset version.

Match the Laboratory Route to the Dosage Question

Platform selection is based on whether markers already exist, how many individuals must be compared, whether subgenome-specific alignment is possible, and how much signal is needed to distinguish neighboring dosage states.

Crop SNP arrays

Best fit: established markers, repeat cohorts, and species where probe behavior and dosage clustering can be calibrated.

Key evidence: raw allele-intensity files, cluster separation, reference samples, and marker annotation—not genotype labels alone.

GBS or reduced-representation sequencing

Best fit: flexible marker discovery across many samples when a fixed array is unavailable or unsuitable.

Key evidence: allele depths, locus consistency, missingness, mapping uniqueness, and population-aware genotype probabilities.

Whole-genome sequencing

Best fit: broad variant discovery, reference improvement, or projects requiring loci outside existing panels.

Key evidence: depth distribution, allele balance, mapping quality, subgenome ambiguity, and validated variant scope.

Targeted confirmation

Best fit: smaller candidate-marker sets, orthogonal checks, or conversion of validated loci into an operational breeding assay.

Key evidence: assay specificity, dosage-cluster behavior, positive controls, and concordance with discovery data.

A route can be staged

A pilot can compare marker recovery and dosage separation before the full cohort is committed. An array can anchor repeat cohorts, GBS can extend marker discovery, WGS can resolve difficult regions or supply reference evidence, and targeted assays can confirm selected loci. The combination is justified only when genome coordinates, alleles, and sample identities can be reconciled.

Available routes include Crop Genotyping Array Services, Genotyping by Sequencing, and Whole Genome Sequencing. Existing crop-specific options include Oat Genotyping Arrays and Wheat Liquid Phase Genotyping Arrays.

Decision routes for polyploid genotyping with SNP arrays, GBS, WGS, and targeted confirmation

How Polyploid Dosage Calling and Quality Control Work

Raw array intensities or sequence allele counts are converted into dosage evidence only after ploidy, locus behavior, sample identity, uncertainty, segregation, and population consistency have been reviewed.

Step 1: Convert Raw Signals into Allele-Dosage Evidence

The workflow is organized around decisions that can stop or redirect the project. Samples and markers do not advance simply because a software command completed.

1. Freeze the biological assumptions

Record expected ploidy, auto- or allopolyploid context, population design, parents, reference assembly, target loci, and downstream data requirements.

2. Review sample and platform evidence

Audit identity, DNA or source-data QC, replicate design, array intensity or allele depth, coverage distribution, missingness, and batch structure.

3. Build the analyzable marker set

Align loci to the agreed reference, remove or flag ambiguous mappings, reconcile alleles, assess homeolog specificity, and retain the signal fields required for dosage estimation.

4. Estimate dosage and uncertainty

Fit the agreed signal or read-count model using suitable ploidy, error, bias, dispersion, frequency, family, or population assumptions; retain posterior evidence where appropriate.

5. Test biological consistency

Evaluate parents and progeny, expected segregation, replicate concordance, ploidy controls, allele-frequency behavior, and inconsistencies that may indicate sample or locus problems.

6. Release for the stated use

Apply project-specific thresholds, compare alternative representations where needed, document exclusions, and create a dosage-aware handoff for the agreed downstream model.

Step 2: Apply Polyploid-Specific Quality Control

A high call rate can hide systematic dosage errors. Quality review must show whether adjacent dosage states separate, whether the expected population supports those states, and whether technical artifacts concentrate in particular samples, loci, subgenomes, or batches.

Polyploid dosage quality gates covering signal separation, depth, segregation, replicates, and homeolog specificity

Five evidence gates before release

  • Signal gate: allele intensity or read-ratio distributions support the proposed dosage states.
  • Confidence gate: depth, posterior probability, genotype quality, and outlier rules separate supported calls from uncertain ones.
  • Population gate: parents, progeny, allele frequencies, and segregation patterns are reviewed against the stated design.
  • Specificity gate: multi-mapping, paralogs, homeologs, probe cross-hybridization, and subgenome assignment are controlled where possible.
  • Reproducibility gate: technical replicates, controls, batches, and platform bridges show acceptable concordance for the intended use.

What happens when the evidence is uneven?

We may retain genotype probabilities instead of hard calls, increase the missing state for ambiguous samples or loci, analyze ploidy groups separately, add parents or reference controls, revise mapping filters, compare a continuous representation, increase data support for a subset, or exclude unsupported regions. The outcome is a documented decision, not a silent conversion of uncertainty into certainty.

Deliverables That Preserve Dosage, Uncertainty, and Provenance

The delivery package is designed so a downstream analyst can identify what was observed, what was inferred, which model and version were used, and where the evidence stops.

Dosage-aware VCF

Ploidy-aware genotype fields with agreed allele-depth, depth, quality, dosage, probability, and filter annotations retained where supported.

Analysis matrices

Discrete dosage, expected dosage, genotype-probability, continuous-ratio, or controlled pseudo-diploid matrices prepared for compatible downstream tools.

Sample and marker QC

Missingness, depth or intensity behavior, confidence, heterozygosity, dosage-state distribution, duplicates, batches, and exclusion decisions.

Population consistency evidence

Parent-offspring checks, segregation summaries, allele-frequency review, ploidy-group behavior, and unresolved biological conflicts.

Marker annotation and provenance

Genome build, coordinates, alleles, subgenome or homeolog status where supported, platform source, caller/model, parameters, and release version.

Downstream readiness record

A clear recommendation for which dataset can enter linkage mapping, association, genomic prediction, diversity analysis, or targeted marker validation.

Validated dosage datasets can support Genetic Linkage Map construction, GWAS Services, or broader Agricultural Genomic Data Analysis. Each downstream use receives the representation and filter boundary it can actually support.

Published Research Case: Dosage-Aware Trait Mapping in Hexaploid Sweet Potato

Haque, E., Shirasawa, K., Suematsu, K., Tabuchi, H., Isobe, S., & Tanaka, M. (2023). Polyploid GWAS reveals the basis of molecular marker development for complex breeding traits including starch content in the storage roots of sweet potato. Frontiers in Plant Science, 14, 1181909. DOI: 10.3389/fpls.2023.1181909.

Research question

Could genome-wide SNP evidence be used in an autohexaploid sweet-potato population to investigate starch content alongside dry matter, storage-root weight, and anthocyanin-related variation?

Study design

The researchers evaluated a 204-member F1 population from two contrasting parents, collected field phenotypes across 2019 and 2020, generated ddRAD sequencing data, retained 90,222 SNPs after alignment and filtering, estimated allele dosage under a six-copy model, and applied polyploid GWAS.

Key findings

The study identified trait-associated signals across the full and phenotype-defined subsets, highlighted a starch-content signal consistent across years, and examined candidate genes in the associated homologous group. The important project lesson is the chain from population design and allele-depth evidence to dosage estimation, trait association, cross-year consistency, and candidate follow-up.

Why it matters for this Solution

  • The biological ploidy model was declared before dosage estimation.
  • Allele-depth evidence was retained rather than reduced immediately to diploid calls.
  • Genotype preparation was connected to a specific downstream polyploid model.
  • Multi-year phenotype evidence helped distinguish persistent from context-specific signals.
  • Association signals were treated as a basis for marker development and biological follow-up, not as finished breeding assays.

What this study does not prove

Its marker effects, thresholds, loci, and performance are specific to the studied cross, reference, sequencing data, traits, environments, and analysis. They do not establish expected call quality or marker transferability for another species, population, platform, or breeding program.

Hexaploid sweet potato case connecting ddRAD allele depths, dosage estimation, field traits, and polyploid GWAS Illustration: original conceptual summary based on the cited study. Not a reproduction of the published figure.

Samples, Data, and Project Information We Can Review

The project can start from existing genotype evidence or from biological material requiring a new data-generation route. Exact sample specifications are confirmed only after the species, tissue, platform, ploidy, and downstream objective are reviewed.

Existing array or sequencing data

  • Raw array intensity files and manifest or probe annotation
  • FASTQ, BAM/CRAM, or VCF with allele-depth and quality fields
  • GBS, ddRAD, WGS, or targeted-genotyping data
  • Previous genotype calls, filters, software, and parameter records

Extracted DNA or biological samples

  • Extracted plant genomic DNA for the selected array or sequencing route
  • Young leaf or other species-appropriate tissue after feasibility confirmation
  • Parents, progeny, ploidy controls, reference accessions, and technical replicates
  • Stable sample identifiers and a complete sample manifest

Biological and downstream context

  • Species, expected ploidy, inheritance context, and known ploidy evidence
  • Population design, pedigree, parentage, and clone or accession relationships
  • Reference assembly, annotation, subgenome information, and marker versions
  • Planned linkage, GWAS, genomic prediction, diversity, or marker-validation use

Information that changes the project route

Tell us whether ploidy varies among samples, whether parent and progeny data are available, whether raw intensity or allele-depth fields have been retained, whether the reference distinguishes subgenomes, and which downstream software or file standard must be supported. These details determine whether a pilot, new genotyping, data rescue, or full-cohort dosage analysis is appropriate.

Why CD Genomics

  • Connected laboratory routes: arrays, GBS, and WGS can be evaluated within one polyploid data plan instead of being treated as interchangeable files.
  • Analysis linked to downstream use: dosage representation and QC are selected around linkage, GWAS, or breeding analysis requirements.
  • Traceable decisions: ploidy assumptions, raw-signal evidence, exclusions, uncertainty, model versions, and handoff boundaries remain documented.

FAQ

1) Why can't standard diploid calling be used for every polyploid dataset?
A diploid call has two homozygous states and one heterozygous state. Polyploids can contain several heterozygous allele-copy states that may carry different inheritance and modeling information. Diploidization is acceptable only when the downstream question deliberately does not require dosage and the information loss is documented.
2) Do we need to know the ploidy before sending samples?
A supported expectation is strongly preferred. Published species knowledge, chromosome counts, flow cytometry, parents, and genome-wide allele ratios can all contribute. If ploidy is unknown or mixed, the project should include a confirmation strategy rather than applying one fixed model to every sample.
3) Can existing array data be rescored for allele dosage?
Potentially, if raw allele-intensity files, array manifests, sample relationships, and suitable control or population information are available. Final genotype labels alone usually do not preserve enough evidence to rebuild dosage clusters reliably.
4) Can GBS data support allele-dosage analysis?
Yes, when allele depths, mapping behavior, locus consistency, missingness, population design, and uncertainty are modeled appropriately. Low or uneven depth may favor genotype probabilities or continuous expected dosages over forced hard calls.
5) How much sequencing depth is required?
There is no universal threshold. Required evidence depends on ploidy, adjacent dosage-state separation, allele and mapping bias, platform, marker number, population priors, and whether the downstream method can use probabilities. A pilot is recommended when the route is uncertain.
6) How are homeologous and paralogous loci handled?
We review mapping uniqueness, reference and subgenome resolution, probe or primer specificity, allele-ratio distortion, excessive depth, unexpected heterozygosity, and population inconsistency. Unsupported loci may be flagged, modeled separately, or excluded; not every region can be assigned unambiguously.
7) Can the delivered matrix be used directly for linkage mapping or GWAS?
It can be prepared for a named compatible method, but the representation must match that method's assumptions. The delivery specifies ploidy, dosage encoding, probabilities or ratios, missing values, filters, population groups, and marker exclusions so the downstream analyst can use the correct version.

Discuss Your Polyploid Genotyping and Dosage Project

Start with the crop, expected ploidy, population design, current data or sample type, and the decision the final genotype matrix must support. We will map those inputs to a pilot, data-rescue, or full-cohort workflow.

  • Species, ploidy evidence, and auto- or allopolyploid context
  • Parents, population structure, pedigree, and target sample count
  • Available array intensity, FASTQ, BAM/CRAM, VCF, or extracted DNA
  • Reference assembly, subgenome resources, and marker annotation
  • Planned linkage map, GWAS, genomic prediction, or marker-validation use

For platform-selection context, see Genotyping Arrays for Polyploid Crops in Wheat, Oat, Brassica, and Cotton.

References

Voorrips, R. E., Gort, G., & Vosman, B. (2011). Genotype calling in tetraploid species from bi-allelic marker data using mixture models. BMC Bioinformatics, 12, 172. DOI: 10.1186/1471-2105-12-172.

Gerard, D., Ferrão, L. F. V., Garcia, A. A. F., & Stephens, M. (2018). Genotyping Polyploids from Messy Sequencing Data. Genetics, 210(3), 789–807. DOI: 10.1534/genetics.118.301468.

Clark, L. V., Lipka, A. E., & Sacks, E. J. (2019). polyRAD: Genotype Calling with Uncertainty from Sequencing Data in Polyploids and Diploids. G3: Genes, Genomes, Genetics, 9(3), 663–673. DOI: 10.1534/g3.118.200913.

Yamamoto, E., Shirasawa, K., Kimura, T., Monden, Y., Tanaka, M., & Isobe, S. (2020). Genetic Mapping in Autohexaploid Sweet Potato with Low-Coverage NGS-Based Genotyping Data. G3: Genes, Genomes, Genetics, 10(8), 2661–2670. DOI: 10.1534/g3.120.401433.

Haque, E., Shirasawa, K., Suematsu, K., Tabuchi, H., Isobe, S., & Tanaka, M. (2023). Polyploid GWAS reveals the basis of molecular marker development for complex breeding traits including starch content in the storage roots of sweet potato. Frontiers in Plant Science, 14, 1181909. DOI: 10.3389/fpls.2023.1181909.

All products and services are For Research Use Only and not for diagnostic or therapeutic use.

For research purposes only, not intended for clinical diagnosis, treatment, or individual health assessments.

Send a MessageSend a Message

For any general inquiries, please fill out the form below.

For research purposes only, not intended for clinical diagnosis, treatment, or individual health assessments.