Genotyping and Genetic Diversity Analysis Services: GBS, ddRAD-Seq, SNP Genotyping, and Population Genomics Solutions
Genotyping — determining the allelic composition of an organism at specific loci — is the foundational measurement of modern genetics. Every population genetics study, every genome-wide association, every breeding program, every conservation management plan begins with the same question: what are the genotypes, and how do they vary across individuals? The methods available to answer that question have expanded dramatically over the past two decades, from gel-based restriction fragment length polymorphism (RFLP) assays that interrogated dozens of loci to next-generation sequencing-based approaches that simultaneously genotype hundreds of thousands of single-nucleotide polymorphisms (SNPs) across hundreds of individuals in a single sequencing run. This article provides a decision-oriented overview of the four dominant genotyping strategies — reduced-representation sequencing (GBS and ddRAD-seq), microsatellite and targeted SNP genotyping, genome-wide association studies, and array-based high-throughput platforms — and offers a framework for selecting the right approach based on species, budget, marker density requirements, and reference genome availability.
Figure
1: Genotyping Technology Evolution — From RFLP to NGS-Based High-Throughput Genotyping
GBS and ddRAD-Seq — Thousands of SNPs Without a Reference Genome
Genotyping-by-sequencing (GBS) and its refined variant double-digest restriction-site-associated DNA sequencing (ddRAD-seq) have become the workhorse methods for population genetics in non-model organisms. The core principle is simple: genomic DNA is digested with one or more restriction enzymes, adapters are ligated to the cut sites, and only the fragments immediately adjacent to restriction sites are sequenced. Because every individual's genome is cut at the same sequence motifs, the same orthologous loci are recovered across all samples — producing a reproducible set of thousands to tens of thousands of SNP markers distributed across the genome.
The method owes its popularity to three features. First, it requires no prior genomic information: no reference genome, no SNP database, no probe design. This makes it immediately applicable to any species, which is why GBS and ddRAD-seq dominate conservation genetics, phylogeography, and the study of wild populations where reference resources are scarce. Second, the per-sample cost is exceptionally low at scale — a 2025 study optimized a ddRAD-based GBS pipeline for up to 384 samples per sequencing run, achieving library preparation costs below 9 euros per sample and sequencing costs of approximately 10 to 14 euros per sample (Fischer et al., BMC Genomics, 2025). Third, the enzyme-based complexity reduction naturally enriches for polymorphic loci: restriction sites that differ between individuals due to SNP or indel variation produce presence/absence patterns in the sequencing data, adding an additional layer of genetic information beyond the SNPs themselves.
The distinction between GBS and ddRAD-seq matters for study design. Standard GBS uses a single restriction enzyme (commonly ApeKI or PstI) and relies on PCR or size selection to reduce fragment complexity. This makes it the most economical option for very large cohorts — breeding panels of thousands of individuals, for example — but introduces variability in which loci are recovered from sample to sample, leading to higher rates of missing data. ddRAD-seq uses two restriction enzymes plus a precise size-selection window (typically 300 to 500 base pairs), which substantially improves locus repeatability across samples and reduces missing data. For population structure analyses where F-statistics, ADMIXTURE, and phylogenetic inference depend on having genotypes at shared loci, ddRAD-seq is generally the safer choice. For breeding programs where the primary output is genomic estimated breeding values from thousands of individuals, the lower per-sample cost of single-enzyme GBS can outweigh its higher missing-data rate, particularly when imputation pipelines are available. For researchers planning reduced-representation projects, genotyping by sequencing and ddRAD-seq services provide end-to-end support from enzyme selection and library preparation through bioinformatic SNP calling and population genetic analysis.
The choice of restriction enzyme deserves more attention than it typically receives. A 2024 study in BMC Genomics analyzed 80 genome assemblies across plants, protostomes, and deuterostomes and found that enzyme recognition site GC content and genome composition significantly bias locus distribution — GC-rich enzyme sites tend to enrich for exonic regions, while AT-rich sites skew toward intergenic and intronic regions (Galla-Camps et al., 2024). The implication is that investigators should consider whether they want markers concentrated in or near genes (useful for adaptation scans and GWAS) or broadly distributed across the genome (better for demographic inference and neutral population structure). In silico digestion tools such as SimRAD and ddgRADer allow researchers to predict fragment distributions for any enzyme combination against a reference genome or a closely related species' assembly before committing to wet-lab work. For a deeper treatment of reduced-representation library preparation, experimental design, and data analysis with Stacks and ipyrad, see our detailed guide on ddRAD-Seq and RAD-Seq for population genetics and phylogenomics.
Figure
2: GBS and ddRAD-Seq Workflow — From Restriction Digestion to SNP Matrix
Microsatellite and SNP Genotyping — Targeted Markers for Parentage, Conservation, and Forensics
While reduced-representation sequencing discovers thousands of anonymous genome-wide SNPs, there are many applications where the question is not "what does the genome look like" but "do these specific known markers match?" Parentage assignment in captive breeding colonies, forensic identification of poached wildlife, cultivar authentication in agriculture, and population monitoring using established marker panels all benefit from genotyping approaches that target specific, well-characterized loci with high information content per marker.
Microsatellites — short tandem repeats of 1 to 6 base-pair motifs that vary in repeat number between individuals — remain widely used in these contexts despite the rise of SNP-based methods. Their multi-allelic nature means a single microsatellite locus carries far more information than a bi-allelic SNP: a locus with 10 alleles can distinguish 55 genotypes, whereas a SNP distinguishes only 3. A panel of 10 to 20 highly polymorphic microsatellites can resolve parentage with near-certainty in managed populations, and because microsatellite genotyping uses standard fragment analysis by capillary electrophoresis — a technology available in virtually every core facility — results are highly portable and reproducible across laboratories and over time. A 2025 validation study demonstrated that 23 microsatellite markers organized into 7 multiplex panels reliably assigned parentage across more than 2,900 macaques from two species, with most markers showing polymorphism information content values above 0.5 and successful amplification from non-invasive samples including hair follicles (de Groot et al., Ecology and Evolution, 2025). Microsatellite genotyping services support project design from marker development and multiplex optimization through fragment analysis and allele scoring for parentage, diversity, and forensic applications.
SNP genotyping of individual markers occupies the middle ground between genome-wide discovery and targeted microsatellite analysis. When a study has already identified the SNPs that matter — from a prior GWAS, from a candidate gene study, or from a GBS discovery run — the most efficient path forward is often to genotype only those validated SNPs across a larger sample set. Platforms for targeted SNP genotyping span a wide range of throughput: TaqMan and KASP assays for one to a few SNPs across thousands of samples; MassARRAY for tens to hundreds of SNPs across hundreds of samples; and targeted amplicon sequencing for hundreds of SNPs across moderate sample numbers. The key advantage over genome-wide methods is cost: when only a small number of markers is needed, paying for whole-genome or reduced-representation sequencing on every sample is wasteful. For a full discussion of marker selection, platform comparison, and study design, see Microsatellite and SNP genotyping services for parentage, conservation, and diversity studies.
Figure
3: Microsatellite vs. SNP Genotyping — Information Content, Throughput, and Application Matching
GWAS — Connecting Genotypes to Phenotypes
Genome-wide association studies (GWAS) transform genotyping data from a catalog of genetic variation into a map of genotype-phenotype relationships. The principle is straightforward: at every marker, the allele is tested for statistical association with the trait of interest, and markers that exceed a genome-wide significance threshold are considered linked to causal variants. In practice, the challenge lies not in running the association tests but in controlling the confounding effects of population structure — the non-random distribution of alleles across subpopulations that can produce spurious associations — and in having sufficient statistical power to detect variants of small effect in traits controlled by many genes.
GBS has become a natural partner for GWAS in non-model organisms because it simultaneously delivers the two things GWAS needs: dense genome-wide markers and genotypes across all individuals in the study. A single GBS run on a diversity panel of 200 individuals typically yields 10,000 to 50,000 SNPs after filtering, which is dense enough to tag most common haplotypes in species with moderate linkage disequilibrium decay distances. The statistical analysis pipeline is well established: SNPs are filtered for minor allele frequency (typically ≥5 percent), call rate (≥80 percent), and Hardy-Weinberg equilibrium, then tested using mixed linear models that incorporate a kinship matrix and principal components to account for population structure. Tools such as GEMMA, GAPIT, and EMMAX implement these models efficiently for GBS-scale datasets, accepting standard VCF genotype files and phenotype tables to produce association statistics that can be visualized with R packages such as qqman for Manhattan and Q-Q plots.
The critical design parameter for GBS-based GWAS is sample size. Power to detect a variant depends on its effect size, its allele frequency, and the number of individuals phenotyped and genotyped. For traits controlled by many loci of small effect — which describes most agronomic and fitness-related traits — sample sizes in the range of 300 to 1,000 individuals are typically needed to detect more than a handful of significant associations. Studies with fewer than 100 individuals are underpowered for all but the largest-effect loci, and even those should be interpreted cautiously due to the winner's curse — the tendency for effect sizes at significant markers to be overestimated in small samples. For researchers planning association mapping studies, genome-wide association study services provide integrated genotyping, phenotyping data management, and statistical analysis pipelines. A comprehensive discussion of experimental design, power analysis, and interpretation is available in our guide on GWAS experimental design with GBS.
Figure
4: GWAS Workflow — From Genotype and Phenotype Data to Manhattan Plot and Candidate Genes
Array-Based High-Throughput Genotyping — Fixed-Content Platforms for Population-Scale Studies
For studies that require consistent, standardized marker sets across thousands of samples — human population genetics consortia, large-scale genomic prediction in agricultural species, biobank-scale cohort studies — SNP microarrays remain the platform of choice. Modern arrays such as the Illumina Infinium Global Diversity Array (GDA) pack approximately 1.8 million markers onto a single bead chip, with content selected to capture common variation across multiple continental populations. Agricultural arrays exist for major crop and livestock species, enabling standardized genomic selection pipelines where genotypes from different studies, laboratories, and countries must be merged and compared.
The core advantage of arrays over sequencing-based genotyping is standardization. Every sample is interrogated at exactly the same set of markers, measured on the same platform, with the same chemistry. There are no missing loci due to stochastic sampling of restriction sites (as in GBS) or variable sequencing depth (as in low-pass WGS). Genotype calling is mature and automated through software such as GenomeStudio, and quality control metrics — call rate, cluster separation, reproducibility — are well characterized and standardized across the field. This makes arrays the default choice for applications where data must be comparable across laboratories and over multi-year study periods: human GWAS consortia, national breeding programs, and population-scale pharmacogenomic research.
The principal limitation of fixed-content arrays is ascertainment bias: the markers on the chip were selected based on patterns of variation in the discovery populations used during array design. For the GDA, this means variation specific to populations underrepresented in early sequencing efforts is systematically missed. Low-pass whole-genome sequencing with imputation is emerging as a competing approach that avoids ascertainment bias entirely — a 2025 comparison across 2,504 individuals from the 1000 Genomes Project found that low-coverage sequencing (0.5x to 2x) performed competitively with population-specific arrays for imputation accuracy and polygenic score estimation, and outperformed arrays in underrepresented populations (Kostyukova et al., 2026). However, for applications where the standard marker set is fit for purpose, arrays remain more economical and computationally tractable at very large scale. For detailed discussion of array platforms, content selection, and comparison with sequencing-based approaches, see our guide on Global Diversity Array and Infinium array genotyping.
Figure
5: Array-Based Genotyping — Infinium BeadChip Technology and Fixed-Content vs. Custom Design
Selecting Your Genotyping Strategy — A Decision Framework
With multiple methods available, each suited to different combinations of species, budget, marker density, and reference resources, the key to efficient genotyping is matching the method to the question. The decision can be organized around four questions.
First: do you have a reference genome? If yes, the full range of methods is available, including imputation-augmented low-pass WGS. If no, reduced-representation methods (GBS and ddRAD-seq) become the default for genome-wide markers, since they can operate de novo using clustering-based pipelines such as Stacks and ipyrad. Microsatellite genotyping also requires no reference genome but needs prior marker development — a one-time investment in library screening and primer design.
Second: how many markers do you need? For parentage assignment and forensic identification, 10 to 30 highly informative microsatellites or SNPs can suffice. For population structure and phylogenetic analysis, 1,000 to 10,000 genome-wide SNPs are typically adequate. For GWAS and genomic prediction, 10,000 to 50,000 markers represent the practical range from GBS; arrays provide 50,000 to 1.8 million depending on the species and platform. For studies requiring maximal marker density or the ability to detect rare variants, whole-genome sequencing — potentially at low coverage with imputation — may be warranted. Whole genome sequencing can serve as both a genotyping platform and a discovery tool for structural variants and rare alleles not captured by reduced-representation or array-based approaches.
Third: how many samples? At small scale (under 50 samples), per-sample library preparation costs dominate and method choice is less constrained by budget — choose the method that gives the best data for the question. At medium scale (50 to 500), GBS and ddRAD-seq become highly cost-effective, and arrays become economical if the species-specific array exists. At large scale (over 500), arrays and highly multiplexed GBS are the dominant options; WGS remains expensive unless very low coverage with imputation is acceptable. Population genomics and evolution studies benefit from careful upfront matching of sample size to marker density, as the statistical power of downstream analyses — F-statistics, ADMIXTURE, phylogenetic inference — depends on both the number of individuals and the number of informative markers.
Fourth: what is the end goal of the analysis? Each method has natural affinities. GBS and ddRAD-seq excel for population structure, phylogeography, and genetic diversity assessment in any species. GWAS pipelines built on GBS data connect genotypes to phenotypes for trait mapping and marker-assisted selection. Microsatellites remain practical for parentage, pedigree reconstruction, and forensic casework where multi-allelic, cross-lab-compatible markers are essential. Arrays dominate human population genetics, genomic selection in major agricultural species, and any application where genotypes from multiple studies must be merged. Bioinformatics services for population genetic analysis support the full spectrum of downstream interpretation, from variant calling and filtering through principal component analysis, phylogenetic tree construction, and selection scan detection.
The genotyping technology landscape in 2026 is defined less by what is technically possible — nearly everything is — and more by the economics of matching method to question at the required scale. The most successful genotyping projects are not those that generate the most data, but those that generate exactly the right data to answer a well-formed question within budget. For any genotyping project, targeted genotyping approaches can be customized to the specific marker set, sample number, and analytical objectives of the study.
Figure
6: Genotyping Strategy Decision Framework — Reference, Markers, Samples, and Goal
FAQ
What is the difference between GBS and ddRAD-seq?
GBS uses a single restriction enzyme to reduce genome complexity, making it the most economical option for very large cohorts but producing higher rates of missing data across samples. ddRAD-seq uses two restriction enzymes plus a precise size-selection window, which improves locus repeatability and reduces missing data — making it the preferred choice when robust population genetic statistics are the priority.
When should I use microsatellites instead of SNPs?
Microsatellites are multi-allelic, so a single locus contains far more information than a bi-allelic SNP — 10 to 20 good microsatellites can resolve parentage with high confidence. They remain the standard for pedigree reconstruction, forensic casework, and any context where results must be cross-comparable across laboratories using standard capillary electrophoresis.
Do I need a reference genome for GBS or ddRAD-seq?
No. Both methods can operate de novo using clustering-based bioinformatic pipelines such as Stacks and ipyrad, which identify orthologous loci across samples without a reference. Having a reference genome improves variant calling and annotation but is not required, which is why these methods dominate research in non-model organisms.
How many samples do I need for a GWAS?
For traits controlled by many loci of small effect, 300 to 1,000 individuals is typically needed to reliably detect significant associations. Studies with fewer than 100 individuals are generally underpowered except for loci with very large effects, and significant hits from small studies should be interpreted cautiously due to the winner's curse.
What are the main limitations of SNP arrays?
The principal limitation is ascertainment bias: the markers on the chip were selected based on variation in the discovery populations used during array design, so variation specific to unrepresented populations is systematically missed. Fixed-content arrays also cannot be updated to include new markers discovered after the array was designed.
How do I choose between reduced-representation sequencing and whole-genome sequencing for population genetics?
If your species has a reference genome and you need the maximum number of markers including rare variants, or if you need structural variant detection, WGS — potentially at low coverage with imputation — is the better choice. If you lack a reference genome, have a limited budget, or need only thousands to tens of thousands of markers for standard population genetic analyses, reduced-representation methods are more practical and cost-effective.
References:
- Fischer D, Tapio M, Bizien T, Sonstebo JH, Runtz T, Hindle MM, Crepaldi P, et al. Fine-tuning GBS data with comparison of reference and mock genome approaches for advancing genomic selection in less studied farmed species. BMC Genomics. 2025;26:111. https://doi.org/10.1186/s12864-025-11296-4
- Kostyukova I, Kenzhebekova R, Turzhanova A, et al. Next-generation genotyping: innovations driving plant genomic improvement. Life. 2026;16(3):521. https://doi.org/10.3390/life16030521
- Galla-Camps M, Carreras C, Pascual M, Pegueroles C. Genome composition and GC content influence loci distribution in reduced representation genomic studies. BMC Genomics. 2024;25:410. https://doi.org/10.1186/s12864-024-10312-3
- de Groot NG, de Vos-Rouweler AJM, Heijmans CMC, Louwerse AL, Massen JJM, Langermans JAM, Bontrop RE, Bruijnesteijn J. Genetic conservation and population management of non-human primates: parentage determination using seven microsatellite-based multiplexes. Ecology and Evolution. 2025;15:e71216. https://doi.org/10.1002/ece3.71216
- de Pontes FCF, Machado IP, Silveira MVS, Lobo ALA, Sabadin F, Fritsche-Neto R, DoVale JC. Combining genotyping approaches improves resolution for association mapping: a case study in tropical maize under water stress conditions. Frontiers in Plant Science. 2025;15:1442008. https://doi.org/10.3389/fpls.2024.1442008
- Li J, Liu Y, Zhang Q, Wang H. Comprehensive review for SSR analysis tools. Frontiers in Genetics. 2024;15:1474611. https://doi.org/10.3389/fgene.2024.1474611
Related Services
For research purposes only, not intended for clinical diagnosis, treatment, or individual health assessments.