Polyploid Genotyping QC: Allele Dosage, Ploidy Errors, Missingness, and Heterozygosity
Many economically important crops—including autotetraploid potato and alfalfa, allohexaploid bread wheat, allooctoploid strawberry, and complex polyploid sugarcane and sweet potato—carry more than two chromosome sets. Unlike diploids, where a biallelic locus is commonly represented by three genotype states, polyploids can contain multiple allele dosage classes. In an autotetraploid, for example, the alternate allele may occur in 0, 1, 2, 3, or 4 copies. Applying diploid assumptions without adjusting for ploidy can collapse intermediate dosage classes, distort heterozygosity estimates, increase genotype uncertainty, and reduce the reliability of downstream association or prediction analyses. This practical guide explains how to evaluate polyploid genotyping quality, when to use discrete dosage calls versus genotype probabilities, how autopolyploid and allopolyploid workflows differ, and which QC signals should be reviewed before linkage mapping, GWAS, or genomic selection.
Key takeaways
- Polyploid QC should preserve allele dosage information rather than forcing every heterozygous genotype into a diploid-style category.
- Autopolyploid and allopolyploid inheritance provide useful starting models, but preferential pairing, segmental allopolyploidy, and locus-specific behavior can produce intermediate patterns.
- There is no universal sequencing-depth threshold for reliable dosage calling; the requirement depends on ploidy, allelic bias, overdispersion, sequencing error, locus quality, and whether hard calls are required.
- Homoeologous and paralogous sequence interference should be evaluated with mapping quality, allele balance, segregation, depth, and subgenome-specific evidence rather than a single filtering rule.
- When dosage uncertainty is substantial, retaining genotype probabilities or expected dosages can be more informative than forcing uncertain markers into discrete states.
For projects requiring dedicated polyploid genotype calling and downstream dosage analysis, see our polyploid genotyping and allele dosage analysis service.
Autopolyploid versus allopolyploid genotyping QC
How do you evaluate polyploid allele dosage accuracy? The first step is to define the biological inheritance model expected for the material being analyzed. Autopolyploids often show polysomic inheritance because multiple homologous chromosome copies can pair during meiosis, whereas established allopolyploids may show predominantly disomic inheritance because pairing occurs preferentially within differentiated subgenomes. However, these categories are not always absolute. Preferential pairing, segmental allopolyploidy, homoeologous exchange, or locus-specific inheritance can cause observed segregation to fall between simple polysomic and disomic expectations.
For standard array QC fundamentals before adding polyploid-specific dosage checks, explore how to read a genotyping array QC report.
| Quality control parameter | Autopolyploid-oriented workflow | Allopolyploid-oriented workflow |
|---|---|---|
| Chromosome pairing model | Often polysomic, with possible multivalent formation and multiple heterozygous dosage classes | Often predominantly disomic within differentiated subgenomes, but preferential or homoeologous pairing should be checked where relevant |
| Genotype representation | Dosage score per locus (0, 1, 2, …, k, where k is ploidy) | Subgenome-specific genotypes where reliable subgenome assignment is possible, or dosage-aware calls when copies cannot be cleanly partitioned |
| Homoeolog versus homolog separation | Primary concern is distinguishing allelic dosage from paralogous mapping and sequencing noise | Additional need to distinguish segregating alleles from fixed differences among homoeologous subgenomes |
| Heterozygosity interpretation | Multiple dosage classes may be biologically valid; double reduction can affect segregation in some loci and crosses | Apparent fixed heterozygosity can reflect subgenome divergence or cross-mapping rather than a segregating heterozygous locus |
| Typical calling strategy | Sequencing read-count models such as updog or polyRAD; array-intensity models such as fitPoly for compatible assay data | Subgenome-aware mapping and calling where possible, combined with dosage-aware or population-aware methods when ambiguity remains |
| Common QC failure mode | Overlapping dosage clusters caused by limited depth, bias, or overdispersion | Cross-mapping among homoeologous regions and incorrect subgenome assignment |
The mathematics of allele dosage calling
In an autotetraploid (4×), each individual carries four chromosome copies, yielding five possible discrete dosage classes for a biallelic locus. If A represents the alternate allele, these states can be written as:
- Nulliplex (0): aaaa
- Simplex (1): Aaaa, with an idealized alternate-read fraction of 0.25
- Duplex (2): AAaa, with an idealized alternate-read fraction of 0.50
- Triplex (3): AAAa, with an idealized alternate-read fraction of 0.75
- Quadruplex (4): AAAA, with an idealized alternate-read fraction of 1.00
The binomial sampling variance challenge
For a simple sequencing model, the number of reads supporting allele A (y) out of total depth (n) can be approximated as y ~ Binomial(n, p), where p = dosage / k. At limited depth, random read sampling can make neighboring dosage states overlap. The simple binomial model is useful for understanding the problem, but real sequencing data frequently depart from it because of sequencing error, allelic bias, overdispersion, PCR effects, and reference-mapping bias. Gerard et al. (2018) explicitly modeled these sources of uncertainty in the polyploid caller updog, showing why read count alone should not be treated as a perfect dosage measurement.
How much sequencing depth is enough?
There is no universal minimum depth for polyploid dosage calling. Required coverage depends on ploidy, locus-specific mapping quality, allelic bias, overdispersion, sequencing error, population information, and the level of genotype certainty required by the downstream analysis. A useful empirical example comes from targeted sequencing of highly heterozygous autotetraploid potato. Uitdewilligen et al. (2013) reported that approximately 60×–80× read depth provided reliable allele copy-number assessment in their targeted sequencing workflow. That range should be interpreted as a study-specific benchmark rather than a general requirement for all tetraploid crops or sequencing designs.
Simulation and probabilistic-genotyping studies show that useful information can also be extracted from substantially lower or uneven sequencing depths when uncertainty is modeled rather than ignored. For shallow GBS or reduced-representation data, the more important question is often not whether every locus can be hard-called, but whether genotype probabilities or expected dosages are sufficiently calibrated for the intended downstream model. Clark et al. (2019) developed polyRAD specifically to retain and propagate genotype uncertainty from low or uneven sequencing depth, while Liao et al. (2021) showed that probabilistic dosages can be used directly in polyploid linkage analysis instead of requiring every uncertain observation to be converted into a discrete genotype.
For strategic comparisons of sequencing depth versus array reliability across breeding populations, consult choosing between LC-WGS, WGS, GBS, and SNP arrays for breeding, or review marker-density considerations in choosing marker density for breeding cohorts.
Hard dosage calls or genotype probabilities?
A central QC decision is whether to release a marker as a discrete dosage state or retain uncertainty. Hard calls are convenient for many linkage, association, and breeding pipelines, but they can create false precision when neighboring dosage clusters overlap. Probabilistic outputs preserve the model's uncertainty and can prevent borderline observations from being treated as equally certain as well-separated calls.
- Use hard dosage calls when: genotype clusters are well separated, replicate concordance is high, parental segregation is consistent, and the downstream software requires integer dosage states.
- Prefer probabilities or expected dosage when: sequencing depth is uneven, allele balance is overdispersed, neighboring dosage states overlap, or the downstream method can incorporate genotype uncertainty.
- Set uncertain calls to missing when: the downstream software cannot use probabilities and validation shows that low-confidence calls create more bias than the loss of marker density.
- Validate thresholds empirically: posterior-probability cutoffs should be calibrated using technical replicates, control crosses, orthogonal genotyping, or downstream sensitivity analysis rather than copied from another crop.
This distinction is important because low-confidence dosage states do not all have the same downstream consequences. A small amount of uncertainty may be tolerable for some genomic prediction workflows, whereas linkage mapping or testing a rare dosage class may require much stronger genotype separation. Liao et al. (2021) provides a practical example of why retaining probabilistic genotypes can improve the use of uncertain polyploid marker data.
Key QC failure modes in polyploid datasets
1. Homoeologous and paralogous sequence interference
In allopolyploids such as bread wheat or cotton, similar sequences across subgenomes can cause short reads to cross-map between homoeologous regions. Paralogous gene families can create a similar problem in both auto- and allopolyploids. The resulting loci may appear persistently heterozygous or show unexpected allele-balance patterns. Suspected homoeologous or paralogous loci should therefore be evaluated using mapping uniqueness, read depth, allele balance, segregation behavior, subgenome-specific coordinates, and population-wide genotype distributions. A marker should not be removed solely because nulliplex or maximum-dosage classes are absent from a particular population.
2. Allelic bias and overdispersion
Restriction-site variation, PCR amplification, capture efficiency, sequencing error, and reference mapping can shift observed allele ratios away from idealized dosage fractions. Gerard et al. (2018) showed that sequencing error, allelic bias, overdispersion, and outlying observations can materially affect polyploid genotype inference. Probabilistic read-count models such as updog explicitly model several of these effects, while polyRAD uses population and linkage information to improve posterior genotype probabilities when read depth is low or uneven.
3. Ploidy errors, aneuploidy, and mixed cytotypes
Breeding collections can contain unexpected ploidy levels, segmental aneuploidy, or mixed cytotypes. Genome-wide allele-balance distributions can flag samples that do not fit the assumed ploidy model, but abnormal cluster patterns are not by themselves a definitive diagnosis of aneuploidy or mixoploidy. Suspect samples should be reviewed together with read depth, chromosome-specific dosage patterns, cytological information when available, and independent sample identity checks.
4. Missingness and dosage uncertainty
Polyploid genotypes should be evaluated as confidence-weighted observations rather than simply present or missing. Posterior-probability and missingness thresholds should be calibrated to the dataset and downstream objective. Ambiguous loci may be retained as genotype probabilities when supported by the analysis software, while poorly reproducible loci can be masked or removed. The appropriate decision depends on replicate concordance, population structure, allele frequency, dosage class balance, and the consequences of a false call for downstream inference.
5. Segregation inconsistency
Control crosses and known pedigrees provide valuable QC information because expected dosage segregation can reveal sample swaps, allele-calling errors, and incorrect inheritance assumptions. However, expected ratios depend on parental dosage, pairing behavior, recombination, and possible double reduction. Deviations from a simple Mendelian expectation should therefore trigger investigation rather than automatic marker removal.
Methodological evidence supports this QC framework directly. In autotetraploid potato, Uitdewilligen et al. (2013) validated allele copy-number calls against KASP genotyping and provided an empirical depth benchmark for their targeted sequencing design. Gerard et al. (2018) demonstrated the importance of modeling sequencing error, allelic bias, and overdispersion in polyploid read-count data. Clark et al. (2019) showed how posterior genotype probabilities can be retained from low or uneven sequencing data, and Liao et al. (2021) evaluated probabilistic dosage information in downstream polyploid linkage analysis.
Choosing QC strategy by data type
The most informative QC signal changes with the genotyping technology. A marker that looks unreliable in one platform may be interpretable in another because arrays, targeted sequencing, GBS, and whole-genome sequencing expose different sources of uncertainty.
| Data type | Primary QC signals | Common polyploid risk | Practical response |
|---|---|---|---|
| SNP array | Allele-intensity cluster separation, replicate concordance, cluster balance | Overlapping dosage clusters or assay cross-hybridization | Use polyploid intensity-clustering methods such as fitPoly and validate ambiguous clusters with controls |
| Targeted sequencing | Per-locus depth, allele balance, mapping quality, replicate agreement | Insufficient read support for neighboring dosage states | Model read-count uncertainty and validate key loci with orthogonal assays |
| GBS / reduced representation | Depth dispersion, missingness, allele bias, restriction-site dropout | Uneven coverage and heterozygote under-calling | Prefer probability-aware callers and avoid universal hard-call thresholds |
| Whole-genome sequencing | Mapping uniqueness, allele balance, local depth, copy-number context | Paralog and homoeolog cross-mapping in repetitive genomes | Use subgenome-aware references or masks where appropriate and retain dosage uncertainty for ambiguous loci |
| Allopolyploid datasets | Subgenome specificity, homoeologous mapping, chromosome-level dosage patterns | Fixed subgenome differences misinterpreted as segregating SNPs | Use subgenome-specific coordinates and segregation evidence before accepting a locus |
| Autopolyploid datasets | Dosage cluster separation, inheritance model, double-reduction sensitivity | Intermediate dosage states compressed or misclassified | Use dosage-aware genotype models and evaluate uncertainty across all dosage classes |
Polyploid genotyping QC & pipeline readiness checklist
- Ploidy Model Verification: Define the expected ploidy and inheritance model for each population, while allowing for preferential pairing, segmental allopolyploidy, or locus-specific departures where evidence supports them.
- Platform-Aware QC: Evaluate the signals appropriate to the technology—array intensity, sequencing depth, allele balance, mapping quality, or posterior genotype probability—rather than applying one threshold to every data type.
- Dosage Estimation Strategy: Match software to the input data. Sequencing read-count data can be analyzed with tools such as updog or polyRAD, while compatible SNP-array intensity data can be clustered with tools such as fitPoly.
- Homoeolog and Paralog Review: Flag suspicious loci using mapping uniqueness, depth, allele ratios, segregation, and subgenome coordinates rather than a single heterozygosity rule.
- Uncertainty Calibration: Define posterior-probability or missingness filters using replicate concordance, control crosses, orthogonal genotyping, or downstream sensitivity analysis.
- Segregation Checks: Compare parental dosage configurations with progeny distributions while accounting for the inheritance model and possible double reduction where biologically relevant.
- Downstream Compatibility: Confirm whether linkage, GWAS, or genomic prediction software accepts integer dosage, expected dosage, or genotype probabilities before finalizing the marker matrix.
For research teams establishing high-density marker sets across polyploid crops, review our genotyping arrays for polyploid crops, crop genotyping array services, and low-coverage whole-genome sequencing (lc-WGS) solutions. Imputation-related QC considerations are discussed in why genotype imputation accuracy fails in breeding populations.
Preparing dosage data for linkage analysis, GWAS, and genomic prediction
A technically clean genotype matrix is not automatically ready for downstream quantitative genetics. The representation of genotype states must match the statistical model. For autopolyploid GWAS, dosage-aware models can test additive and alternative gene-action assumptions instead of collapsing all heterozygous states into one class. Rosyara et al. (2016) developed GWASpoly to model additive, simplex-dominant, and duplex-dominant effects in autopolyploids and demonstrated the framework in tetraploid potato.
For linkage analysis, uncertain dosage states may be propagated as probabilities when the method supports them, rather than discarded or forced into discrete calls. For genomic prediction, expected dosage values can often be incorporated into relationship matrices, but the treatment of uncertainty should be documented and validated because systematic dosage bias can alter genetic similarity estimates. Before release, teams should record the ploidy model, software version, genotype representation, filtering logic, excluded loci, uncertainty metrics, and any orthogonal validation results so that later analyses remain reproducible.
Standardized marker formatting and downstream QC are further discussed in building GS-ready datasets from array and sequencing outputs, while broader quantitative genetics principles are covered in genomic selection in plant and animal breeding.
How CD Genomics can help
Genotyping and analyzing polyploid crop genomes requires platform-aware laboratory design, dosage-conscious bioinformatics, and quantitative genetics workflows that match the organism's inheritance model. CD Genomics supports polyploid research projects with SNP array genotyping, sequencing-based marker generation, probabilistic allele dosage analysis, subgenome-aware QC, and downstream association or genomic prediction workflows. Our polyploid genotyping and allele dosage analysis service can be integrated with agricultural genomic data analysis and broader molecular breeding and genotyping workflows. Deliverables can include dosage matrices, confidence metrics, QC summaries, marker-level annotations, and analysis-ready files with the assumptions and filtering steps documented for reproducibility. All laboratory and analytical services described herein are provided strictly for Research Use Only (RUO) in agricultural research and breeding programs and are not intended for clinical or diagnostic use.
Frequently asked questions (FAQ)
Q1: Can general-purpose variant callers be used for polyploid samples?
A: Some general-purpose callers, including GATK HaplotypeCaller, can be configured for non-diploid sample ploidy. However, setting ploidy alone does not solve every polyploid genotyping problem. Reliable dosage interpretation still depends on read depth, allele balance, sequencing error, mapping ambiguity, genotype uncertainty, and the organism's inheritance model. Dedicated tools such as updog or polyRAD can explicitly model several of these sources of dosage uncertainty.
Q2: What is the minimum sequencing depth needed for reliable tetraploid allele dosage calling?
A: There is no universal minimum. In one targeted-sequencing study of autotetraploid potato, approximately 60×–80× was reported as a lower boundary for reliable allele copy-number assessment in that specific workflow. Other methods can extract useful dosage information at lower depth by modeling genotype uncertainty. The appropriate depth should therefore be validated for the crop, platform, locus complexity, and downstream requirement for hard calls versus probabilistic dosage.
Q3: How do I distinguish a true segregating SNP from a homoeologous sequence variant in an allopolyploid?
A: Do not rely on a single heterozygosity rule. Evaluate mapping uniqueness, subgenome-specific coordinates, read depth, allele-balance consistency, population-wide dosage patterns, and segregation in known crosses. A fixed difference between subgenomes may create persistent apparent heterozygosity when homoeologous reads are co-mapped, whereas a true segregating allele should show population or pedigree variation consistent with the underlying inheritance model.
Q4: What is double reduction and how does it affect autopolyploid QC?
A: Double reduction occurs when sister chromatid segments can enter the same gamete following multivalent pairing and recombination. It can change expected segregation and identity-by-descent patterns, so linkage or pedigree models may need to account for it when the species, chromosome region, and cross design make double reduction biologically relevant.
Q5: Can polyploid dosage matrices be analyzed with standard diploid GWAS workflows?
A: Standard diploid workflows generally do not natively model the full 0–k allele-dosage range or polyploid-specific gene action. For autopolyploid association studies, dosage-aware tools such as GWASpoly can model additive, simplex-dominant, and duplex-dominant effects. Other software may accept continuous dosage or appropriately encoded covariates, but the representation should be checked against the statistical assumptions before analysis.
References
- Uitdewilligen, Jan G. A. M. L., Anne-Marie A. Wolters, Bjorn B. D'hoop, Theo J. A. Borm, Richard G. F. Visser, and Herman J. van Eck. "A Next-Generation Sequencing Method for Genotyping-by-Sequencing of Highly Heterozygous Autotetraploid Potato." PLOS ONE, vol. 8, no. 5, 2013, e62355.
- Gerard, David, Luís Felipe Ventorim Ferrão, Antonio Augusto Franco Garcia, and Matthew Stephens. "Genotyping Polyploids from Messy Sequencing Data." Genetics, vol. 210, no. 3, 2018, pp. 789–807.
- Clark, Lindsay V., Alexander E. Lipka, and Erik J. Sacks. "polyRAD: Genotype Calling with Uncertainty from Sequencing Data in Polyploids and Diploids." G3: Genes, Genomes, Genetics, vol. 9, no. 3, 2019, pp. 663–673.
- Liao, Yanlin, Roeland E. Voorrips, Peter M. Bourke, Giorgio Tumino, Paul Arens, Richard G. F. Visser, Marinus J. M. Smulders, and Chris Maliepaard. "Using Probabilistic Genotypes in Linkage Analysis of Polyploids." Theoretical and Applied Genetics, vol. 134, no. 8, 2021, pp. 2443–2457.
- Rosyara, Umesh R., Walter S. De Jong, David S. Douches, and Jeffrey B. Endelman. "Software for Genome-Wide Association Studies in Autopolyploids and Its Application to Potato." The Plant Genome, vol. 9, no. 2, 2016.
Send a MessageFor any general inquiries, please fill out the form below.


