Targeted Resequencing vs Whole-Genome Sequencing for Large Cohorts: When Are Candidate Regions Enough?

When managing population genomics projects involving hundreds or thousands of samples, researchers face a critical strategic fork: should you sequence the complete genome across all individuals using whole-genome sequencing (WGS), or concentrate your sequencing budget on pre-defined candidate regions via targeted resequencing?
When is targeted resequencing enough? Targeted resequencing is sufficient when candidate genomic regions, validated disease loci, or quantitative trait loci (QTL) intervals are already identified, target panel sizes remain under ~10–50 Mb, and cohort scale reaches hundreds to thousands of samples with budget constraints. By allocating sequencing read throughput strictly to targeted loci, hybrid capture and amplicon panels achieve high coverage depth (100×–500×+) at a fraction of the per-sample cost of WGS, providing superior statistical power for rare variant detection and accurate genotype calling. Conversely, WGS remains indispensable when novel variant discovery across non-coding regions, unanchored structural variant profiling, or genome-wide haplotype phasing is required.
You will find a decision framework, a sample size by target footprint scenario matrix, a cost driver comparison table, and a Go/Adjust/Stop pilot scorecard to validate your project before cohort expansion.
TL;DR
- Targeted resequencing often wins when candidate loci are defined, sample size (N) is high, and deep coverage is required at predictable per-sample costs.
- Whole-genome re-sequencing wins when unconstrained discovery across non-coding space, structural variation, or genome-wide LD profiling is required.
- Run a pilot study first to evaluate capture efficiency, fold-80 penalty, and duplicate rates using our Go/Adjust/Stop scorecard.
Decision Framework: Targeted Resequencing vs Whole-Genome Sequencing
Choosing between targeted resequencing and whole-genome sequencing (WGS) for large population cohorts requires balancing target region footprint, sample size (N), required coverage depth, and total available capital.
1.1 Direct Answer: When Targeted Resequencing is Sufficient
Targeted resequencing represents the optimal strategy under the following project parameters:
- Pre-existing genomic hypotheses: Prior genome-wide association studies (GWAS), QTL mapping, or transcriptomic profiling have narrowed the regions of interest to specific genes, pathways, or linkage disequilibrium (LD) blocks.
- Large cohort size (N ≥ 500): The study aims to screen large numbers of biological or clinical samples where per-sample sequencing expenditure must be minimized without compromising diagnostic sensitivity.
- High coverage depth requirements: Detecting ultra-rare somatic mutations, mosaic variants, or low-frequency germline alleles requires deep fold coverage (e.g., >100× to 500×+) that would be financially prohibitive across whole genomes.
- Routine or standardized screening: The study serves as a screening phase in molecular breeding, population surveillance, or clinical cohort validation where standardized panel content ensures uniform cross-batch comparisons.
For detailed operational guidance on targeted panel options, consult our Targeted Resequencing Service.
1.2 When Whole-Genome Sequencing Remains Indispensable
Whole-genome sequencing remains the gold standard and necessary choice when:
- Hypothesis-free discovery: The genetic architecture of the phenotype is entirely unknown, requiring comprehensive scanning across coding and non-coding regions.
- Structural variation and complex rearrangements: The study focuses on large copy-number variations (CNVs), balanced translocations, retrotransposon insertions, or complex inversions that span beyond probe capture boundaries.
- Unbiased haplotype phasing and genome-wide LD: The project demands genome-wide imputation backbones or fine-scale population structure profiling across unconstrained intergenic regions.
When comprehensive genomic coverage is required, refer to our Whole Genome Re-sequencing Service.
1.3 Key Trade-offs at a Glance
| Feature / Metric | Targeted Resequencing | Whole-Genome Sequencing (WGS) |
| Primary Focus | Deep interrogation of defined genomic intervals | Unbiased, comprehensive genome-wide analysis |
| Typical Target Footprint | 10 kb to 50 Mb (custom panel or exome) | Entire genome (~3.2 Gb in human) |
| Achievable Coverage Depth | High to ultra-high (100×–1,000×+) | Moderate (15×–30× standard, 60× deep) |
| Cost per Sample (N > 1,000) | Low to moderate ($20–$150) | High ($300–$800+) |
| Rare Variant Sensitivity (<1% MAF) | Superior within target loci due to high depth | Dependent on depth; limited at 15×–30× |
| Non-coding / Intergenic Coverage | Minimal (restricted to targeted flanking regions) | Complete genome-wide coverage |
| Bioinformatics / Storage Overhead | Small data footprint (~0.5–2 GB BAM per sample) | Massive data footprint (~30–100 GB BAM per sample) |
Scenario Matrix: Balancing Sample Size, Target Size, and Coverage Depth
Determining whether targeted resequencing or WGS yields higher return on investment requires evaluating the interaction between target panel footprint (in megabases, Mb), cohort sample size (N), and target depth.
Figure 2. Sample size by target panel footprint and depth matrix for choosing targeted resequencing or WGS.
2.1 The Sample Size × Target Footprint Trade-off
The financial break-even point between targeted capture and WGS shifts as a function of panel size. For small target footprints (<5 Mb), library capture probes account for a minor cost component relative to sequencing reads, making targeted capture exponentially cheaper per sample. As panel size expands beyond 50 Mb (approaching whole-exome scale), the combined cost of custom capture probes, bait synthesis, and sequencing approaches the baseline cost of shallow or standard WGS.
2.2 Coverage Depth Dynamics: High-Fold Capture vs Shallow WGS
In large population cohorts, statistical power for association testing depends on both sample size and genotype accuracy. High-fold targeted resequencing (100×+) virtually eliminates false-positive variant calls and heterozygous dropout within candidate loci. In contrast, low-pass WGS (0.5×–5×) relies on genotype likelihoods and imputation against reference panels, which may introduce uncertainty for rare variants or underrepresented populations. For details on whole-exome strategies, explore our Whole Exome Sequencing Service. For additional perspectives on array vs low-pass vs deep WGS trade-offs, see Arrays vs Low-Pass vs Deep WGS.
2.3 Scenario Matrix for Large Cohort Planning
| Scenario Category | Cohort Size (N) | Target Footprint | Recommended Depth | Recommended Strategy | Primary Rationale |
| Narrow Candidate Loci | 1,000–10,000+ | <2 Mb | 100×–300× | Custom Hybrid Capture Panel | Maximum sample throughput at minimal cost; high rare variant power |
| Multi-Gene Pathways / QTLs | 500–2,000 | 2–20 Mb | 50×–100× | Targeted Resequencing | Balanced trade-off between target depth and cohort scalability |
| Broad Exome Capture | 200–1,000 | 30–50 Mb | 30×–50× | Whole Exome Sequencing (WES) | Focuses on protein-coding variation across all genes; lower cost than WGS |
| Large-Scale Biobank Screening | 5,000–50,000+ | Pre-selected SNPs | 20×–50× (Targeted) | Targeted SNP Panel / Genotyping | Ultra-low cost per sample for rapid population-level screening |
| Unconstrained Discovery | 100–500 | Whole Genome (~3 Gb) | 30× | Whole-Genome Sequencing (WGS) | Unbiased discovery across non-coding, regulatory, and structural domains |
For comparing panel capture vs PCR-based approaches, see our technical comparison Hybrid Capture vs Amplicon Sequencing.
Biological & Variant Discovery Trade-offs in Large Cohorts
Beyond raw cost calculations, researchers must evaluate how each technology impacts biological discovery, variant sensitivity, and analytical accuracy.
3.1 Rare Variant Sensitivity and Statistical Power
Rare variants (minor allele frequency, MAF < 1%) account for a substantial proportion of missing heritability in complex traits and disease susceptibility. Detecting rare variants in population cohorts requires high read coverage to distinguish true biological single nucleotide polymorphisms (SNPs) and insertions/deletions (indels) from sequencing noise.
- Targeted Resequencing: High coverage depth (100×+) ensures that rare alleles are backed by multiple independent reads across both strands, enabling robust caller performance without heavy reliance on population-level imputation.
- WGS at 30×: While sufficient for common and low-frequency variants, 30× WGS may suffer from stochastic sampling loss at rare heterozygous sites, requiring larger sample sizes to reach equivalent call confidence.
When high-density marker screening is sufficient for known loci, review our SNP Genotyping Service.
3.2 Haplotype Resolution, Phase Capture, and Regulatory Context
Targeted capture protocols successfully isolate specific coding and promoter regions. However, physical capture boundaries can truncate long-range linkage disequilibrium (LD) structures and distant enhancer-promoter interactions. If your research hypothesis relies on identifying long-range haplotypes or non-coding regulatory variants located hundreds of kilobases away from the primary gene body, WGS or supplementary long-read sequencing provides contiguous context that targeted panels cannot capture.
3.3 Limitations in Non-Coding Regions and Novel Locus Discovery
Targeted resequencing is inherently hypothesis-driven. If a causal mutation lies in an unannotated intergenic enhancer, deep intronic splice-altering position, or structural expansion outside the probe capture set, targeted resequencing will fail to detect it. WGS avoids capture bias, providing an immutable genomic baseline that can be re-analyzed in silico as functional annotations evolve.
Economic & Budget Structure Analysis for Large-Scale Cohorts
Budget planning for large-scale population genomics extends beyond flow cell reagent costs. A total cost of ownership (TCO) evaluation must account for upfront panel design, library preparation, sequencing, computational infrastructure, and long-term storage.
Figure 3. Cost driver breakdown showing where targeted resequencing achieves maximum budget efficiency.
4.1 Cost per Sample vs Fixed Assay Setup Costs
- Targeted Resequencing: Involves a fixed upfront cost for custom probe synthesis or multiplex primer design. However, as sample size (N) scales from hundreds to thousands, the initial probe synthesis cost is amortized across the cohort, drastically reducing the marginal cost per sample.
- Whole-Genome Sequencing: Features minimal upfront design costs but incurs high, linear reagent and sequencing costs per sample that do not decrease significantly with sample volume.
4.2 Compute, Storage, and Downstream Bioinformatic Overhead
A major hidden expenditure in population genomics is high-performance computing (HPC) and cloud storage:
- WGS Data Footprint: A cohort of 1,000 WGS samples at 30× depth generates approximately 100 TB of raw FASTQ and aligned BAM files, incurring substantial ongoing cloud storage and re-alignment costs.
- Targeted Resequencing Footprint: The same cohort of 1,000 targeted samples at 100× depth generates less than 2–5 TB of total alignment data, reducing computational processing time from days to hours and lowering storage costs by over 90%.
4.3 Cost Driver Breakdown Comparison
| Cost Component | Targeted Resequencing (Custom Panel) | Whole-Genome Sequencing (30× WGS) | Scalability Impact |
| Panel Synthesis / Design | Fixed initial fee ($1,000–$5,000) | $0 (No panel design needed) | Amortized across N; negligible for N > 1,000 |
| Library Preparation | Moderate (Enrichment capture / PCR) | Standard automated library prep | Comparable at scale with automation |
| Sequencing Reagents | Low ($15–$80 per sample) | High ($300–$600 per sample) | Major cost differentiator at scale |
| Primary Bioinformatic Alignment | Fast (~15–30 CPU minutes/sample) | Intensive (~10–20 CPU hours/sample) | 95% compute reduction for targeted data |
| Data Storage (BAM/CRAM + VCF) | ~1–3 GB per sample | ~30–60 GB per sample | 90–95% cloud storage savings for targeted data |
For operational frameworks on managing large-scale cohorts, see Scaling Targeted Resequencing to Large Cohorts.
Operational Roadmap: Transitioning Loci to a Scalable Cohort Panel
Successfully deploying a targeted resequencing panel for large population cohorts requires a structured pipeline from locus definition to quality control.
5.1 Defining Loci Footprints from Prior GWAS, QTL, or Literature
- Collate Lead Signals: Extract significant lead SNPs or QTL peak intervals from discovery GWAS or literature.
- Define Linkage Disequilibrium (LD) Windows: Expand target boundaries around lead markers to encompass surrounding LD blocks (r2 ≥ 0.8) to ensure all potential causal variants are captured.
- Include Regulatory Annotations: Incorporate promoter regions, known eQTLs, and conserved non-coding elements flanking the target gene bodies.
5.2 Hybrid Capture Probe Design, GC Uniformity, and Off-Target Minimization
- Mask Repeat Elements: Filter out low-complexity sequences, retrotransposons, and highly repetitive genomic regions during probe selection to minimize non-specific binding.
- Balance GC Content: Adjust probe density and tiling overlap across high-GC (e.g., CpG islands, promoter regions) and low-GC regions to prevent coverage dropouts.
- Optimize Bait Tiling: Utilize 2× or 3× flexible bait tiling across critical exons and active sites to maintain uniform fold coverage.
5.3 Batch Effects, Cross-Run Consistency, and Quality Control Guidelines
In multi-center or longitudinal cohort studies, batch-to-batch variation can introduce systematic biases. Key mitigation practices include:
- Inclusion of Internal Controls: Include bridge samples (standardized reference DNA) across every sequencing run to monitor capture efficiency and allele frequency consistency.
- Standardized QC Metrics: Enforce strict pass/fail thresholds for target specificity (% reads on target > 70%), fold-80 penalty (< 1.5), and duplicate rates (< 15%).
Pilot Validation & Go / Adjust / Stop Scorecard
Before scaling a targeted resequencing panel to thousands of cohort samples, executing a pilot study (24–48 representative samples) is essential to validate capture performance, library complexity, and target uniformity.
Figure 4. Pilot scorecard for evaluating capture efficiency, duplicate rate, and coverage uniformity before scaling.
6.1 Purpose of a Small Pilot Study
A pilot project provides empirical data on capture efficiency, off-target binding, probe competition, and library duplicate rates. Identifying underperforming probes or high duplicate rates during the pilot stage allows for probe re-balancing or protocol adjustments before committing major capital.
6.2 Key Pilot Readout Metrics
Evaluate the following primary performance indicators during pilot review:
- Mean Target Coverage Depth: Verifies that sequencing throughput meets planned depth targets.
- Percentage of Target Bases at ≥ 30× Coverage: Measures coverage completeness across all targeted intervals.
- Fold-80 Penalty: Quantifies coverage uniformity (a value closer to 1.0 indicates perfect uniformity across targets).
- On-Target Read Ratio: Percentage of total mapped reads aligning directly to target probe coordinates.
6.3 Go / Adjust / Stop Scorecard
| Performance Metric | Go (Proceed to Scale) | Adjust (Optimize Assay) | Stop (Re-evaluate Strategy) |
| On-Target Read Ratio | ≥ 70% on target | 50% - 69% (Adjust hybridization/wash temp) | <50% (High non-specific binding; redesign) |
| Bases at ≥ 30× Depth | ≥ 95% of target footprint | 85% - 94% (Increase depth per sample) | <85% (Severe probe dropout or GC bias) |
| Fold-80 Penalty | ≤ 1.6 (High uniformity) | 1.7 - 2.0 (Re-balance bait concentration) | >2.0 (Extreme unevenness across targets) |
| Library Duplicate Rate | < 15% duplicates | 15% - 25% (Increase initial DNA input) | >25% (Low library complexity) |
| Cross-Batch Concordance | ≥ 99.5% genotype match | 98.0% - 99.4% (Normalize pipeline callers) | <98.0% (Uncontrolled batch effect across runs) |
For comprehensive study planning guidelines, see our guide Population Genomics Study Design Guide.
FAQs
Custom targeted resequencing via hybrid capture is highly efficient for target footprints ranging from a few kilobases up to 20–30 Mb. When panel footprint expands beyond 40–50 Mb, whole-exome sequencing (WES) or standard whole-genome sequencing (WGS) often becomes more cost-effective due to decreasing bait synthesis efficiency and higher probe costs.
Targeted resequencing delivers significantly higher sequencing depth (often 100×–500×+) at targeted candidate loci compared to standard 30× WGS for the same cost per sample. This elevated coverage provides superior statistical confidence and higher call rate accuracy for detecting ultra-rare variants (MAF < 1%) within the targeted candidate regions.
Yes, targeted resequencing data can be harmonized with WGS datasets at overlapping genomic coordinates, provided that variant calling pipelines follow standardized functional equivalence guidelines and reference genome builds (e.g., GRCh38). Joint variant calling across overlapping loci allows cohort data to be merged seamlessly.
Designing a custom hybrid capture panel typically becomes economically advantageous when cohort size reaches 100 to 200 or more samples. At this scale, the fixed upfront probe design and bait synthesis fees are amortized, resulting in a substantially lower total cost per sample compared to running 30× WGS across the entire cohort.
Low on-target ratios are usually caused by incomplete blocking of repetitive genomic sequences, suboptimal hybridization temperatures, or excessive PCR amplification cycles. This issue can be resolved by optimizing repeat-masking during probe design, utilizing high-stringency blocking oligos, and reducing library amplification cycles prior to capture.
Next steps: If you are planning a large cohort study and want to evaluate custom panel feasibility or pilot design, you can discuss your project requirements with our technical team.
Planning note: Numerical ranges, sequencing depths, data-volume estimates, cost estimates, pilot sizes, and QC thresholds presented in this article are synthesized from published literature and publicly reported industry practices and are provided for research planning reference only. Actual project specifications, performance metrics, costs, and acceptance criteria may vary substantially with species, genome size, sample quality, target design, sequencing platform, laboratory workflow, cohort size, and study objectives. Project-specific parameters should therefore be confirmed during experimental design and pilot evaluation.
References:
- Mertes, F., ElSharawy, A., Sauer, S., et al. "Targeted enrichment of genomic DNA regions for next-generation sequencing." Briefings in Functional Genomics, 2011, 10(6): 374–386. doi:10.1093/bfgp/elr033.
- DePristo, M. A., Banks, E., Poplin, R., et al. "A framework for variation discovery and genotyping using next-generation DNA sequencing data." Nature Genetics, 2011, 43(5): 491–498. doi:10.1038/ng.806.
- Indap, A. R., Cole, R., Runge, C. L., Marth, G. T., and Olivier, M. "Variant discovery in targeted resequencing using whole genome amplified DNA." BMC Genomics, 2013, 14: 468. doi:10.1186/1471-2164-14-468.
- Regier, A. A., Farjoun, Y., Larson, D. E., et al. "Functional equivalence of genome sequencing analysis pipelines enables harmonized variant calling across human genetics projects." Nature Communications, 2018, 9(1): 4038. doi:10.1038/s41467-018-06159-4.
- Gnirke, A., Melnikov, A., Maguire, J., et al. "Solution hybrid selection with ultra-long oligonucleotides for massively parallel targeted sequencing." Nature Biotechnology, 2009, 27: 182–189. doi:10.1038/nbt.1523.
- Li, B., and Leal, S. M. "Methods for detecting associations with rare variants for common diseases: application to analysis of sequence data." The American Journal of Human Genetics, 2008, 83(3): 311–321. doi:10.1016/j.ajhg.2008.06.024.
- Besser, J., Carleton, H. A., Gerner-Smidt, P., Lindsey, R. L., and Trees, E. "Next-generation sequencing technologies and their application to the study and control of bacterial infections." Clinical Microbiology and Infection, 2018, 24(4): 335–341. doi:10.1016/j.cmi.2017.10.013.