How to Define Target Regions for Population-Scale Targeted Resequencing: Genes, GWAS Loci, QTL Intervals, or Regulatory Regions?
Figure 1. Each evidence type requires a different rule for converting a biological hypothesis into genomic coordinates.
A targeted resequencing panel is only as clear as its region list. "Sequence these genes" may mean coding exons, untranslated regions, promoters, full gene bodies, or every interval linked to those genes. "Cover this GWAS locus" may refer to the lead SNP, its credible set, an LD block, or a fixed window. QTL coordinates can span millions of bases, and regulatory annotations can multiply faster than the panel budget allows.
The design task is therefore not to collect every plausible interval. It is to translate prior evidence into a versioned set of coordinates that supports a defined analysis at the required depth across the intended cohort. The Targeted Resequencing service can support project-specific enrichment and sequencing, but researchers should first decide what evidence qualifies a region for inclusion and what variants the analysis must be able to observe.
TL;DR
- Candidate genes, GWAS loci, QTL intervals, and regulatory regions need separate coordinate rules.
- A gene symbol is not a target definition. Specify transcript policy, exon treatment, untranslated regions, splice boundaries, and flanking sequence.
- A lead GWAS SNP rarely represents the full locus. Use LD, fine mapping, ancestry, and functional evidence to define the interval.
- A broad QTL should be narrowed before panel design whenever possible.
- Regulatory regions should be included only when the tissue, cell type, species, and evidence match the research question.
- Balance total target size, expected on-target performance, required depth, sample number, and analysis plan together.
Begin with the Intended Inference
Before drawing intervals, write one sentence describing what the sequencing data must answer. Examples include identifying rare coding variants in established genes, resolving variants around replicated GWAS signals, comparing haplotypes within narrowed QTL intervals, or testing variants in experimentally supported regulatory elements. Each statement points to different targets and different analytical units.
Then identify the variant classes of interest:
- Single-nucleotide variants only, or small insertions and deletions as well?
- Coding changes, splice-region variants, untranslated regions, or noncoding variants?
- Common variants, rare variants, or both?
- Known variants, novel discovery within the targets, or replication of specific alleles?
- Single-marker tests, gene-based aggregation, haplotype analysis, or functional prioritization?
These choices influence target breadth, read depth, bait or primer design, and downstream filtering. They also determine whether variant calling, phasing, or region-based association testing must be included in the analysis scope.
Four Evidence Types, Four Definition Rules
Figure 2. The evidence source determines whether the design should follow transcripts, LD, recombination boundaries, or functional annotations.
| Evidence source | Minimum defensible target | Common expansion | Main risk |
| Candidate gene | Defined transcript exons plus splice boundaries | UTRs, promoter, conserved elements, or full gene body | Ambiguous transcripts and missed noncoding regions |
| GWAS locus | Lead SNP and credible or LD-supported variants | LD block, independent signals, nearby functional elements | Arbitrary fixed windows or ancestry mismatch |
| QTL interval | Flanking-marker interval on a named assembly | Fine-mapped subintervals and candidate genes | Panel dominated by an interval that remains too broad |
| Regulatory region | Experimentally or computationally supported element in relevant tissue | Linked promoter, enhancer, accessibility peak, or conserved element | Annotation overload and weak target-to-gene evidence |
Candidate Genes: Define the Transcript Policy
For candidate genes, decide whether the panel targets all annotated transcripts, one canonical transcript, or a project-specific isoform set. Exon coordinates can differ substantially across transcripts. A coding-only design should state whether untranslated exons are excluded and how splice junctions are handled. If the objective includes regulatory or structural variation, exon-only capture may be insufficient.
Useful gene-level options include:
- All coding exons for the selected transcript set.
- Coding exons plus short intronic flanks around splice junctions.
- Coding exons and 5′/3′ untranslated regions.
- Full gene body from transcription start to transcription end.
- Gene body plus a defined upstream promoter window.
- Selected conserved or experimentally supported elements linked to the gene.
Do not apply one expansion rule to every gene without considering size and evidence. A full-gene design may be practical for compact genes but consume a large fraction of the panel for genes with long introns. Record the annotation release and transcript identifiers so the target set can be reproduced.
GWAS Loci: Move Beyond the Lead SNP
A lead SNP is a statistical marker, not automatically the causal allele or the boundary of the locus. Target definition should consider conditionally independent signals, the credible set from fine mapping, population-specific LD, recombination hotspots, nearby genes, and functional annotations. The linkage disequilibrium overview explains why the same lead signal can correspond to different correlated variants across ancestries.
A practical hierarchy is:
- Include the lead SNP and every variant in a stated credible set, if available.
- Add variants in strong LD using the study population or a justified reference population.
- Represent independent signals separately rather than merging them into one undifferentiated interval.
- Add coding or regulatory elements only when there is evidence linking them to the association.
- Retain enough flanking sequence for enrichment design and reliable alignment.
Fixed windows can be used when detailed fine mapping is unavailable, but the window size must be justified and should not be presented as a biological boundary. For multi-ancestry projects, combine or compare ancestry-specific credible sets rather than using one LD block for every group. Statistical fine mapping can reduce the region, but it does not prove causality.
QTL Intervals: Narrow Before You Tile
QTL intervals are often defined by flanking markers or confidence intervals and may cover many genes. Directly targeting an entire broad interval can make a "targeted" panel approach whole-genome scale while still missing important sequence outside the chosen boundaries. Use recombinants, additional markers, multiple populations, expression data, comparative genomics, or prior fine mapping to narrow the interval first.
The design file should state:
- Trait and mapping population.
- QTL name and the exact flanking markers.
- Genetic and physical interval, including assembly version.
- Confidence interval or support interval method.
- Evidence that the QTL replicates across environments or backgrounds.
- Candidate genes or subregions selected from the broader interval.
- Whether structural variants or repetitive sequence are expected.
The comparison of QTL and GWAS study designs is useful when the project combines family-based linkage evidence with population association evidence. When both converge on the same interval, that convergence can justify higher priority, but target boundaries still need explicit coordinates.
Regulatory Regions: Require Contextual Evidence
Promoters and enhancers are not interchangeable across tissues, developmental stages, or environmental conditions. Regulatory targets should have a documented source: accessible chromatin, histone marks, transcription-factor binding, expression QTL evidence, reporter assays, conservation, or another relevant dataset. State the reference assembly and annotation release.
For promoters, define the anchor and window, such as a stated distance upstream and downstream of a transcript start site. For enhancers, use the experiment-derived peak or curated interval rather than a vague distance from the gene. When enhancer-to-gene linkage is uncertain, label it as a hypothesis rather than a confirmed relationship. Avoid adding every predicted regulatory element near each candidate gene; that approach expands cost without a corresponding gain in interpretability.
Information Needed to Design Target Regions
| Input | Required detail | Example decision enabled |
| Reference assembly | Species, assembly name, patch/version | Coordinate normalization and bait design |
| Annotation release | Gene model source and version | Transcript and exon selection |
| Evidence class | Gene, GWAS, QTL, regulatory, or fixed variant | Applies the correct inclusion rule |
| Coordinates | Chromosome, start, end, strand where relevant | Creates a BED-like target file |
| Evidence source | DOI, database record, analysis result, or experiment | Supports prioritization and auditability |
| Population | Ancestry, breed, stock, or mapping population | Guides LD, frequency, and reference choices |
| Variant classes | SNV, indel, structural, repeat-associated | Determines method feasibility |
| Desired depth | Minimum usable depth and uniformity goal | Links panel size to sequencing allocation |
| Cohort size | Samples, groups, batches, and replicates | Determines multiplexing and total data need |
| Analysis endpoint | Calling, association, haplotypes, burden test, annotation | Defines required coverage and deliverables |
| Priority | Required, high, medium, optional | Supports trimming when design constraints arise |
The coordinate file should be accompanied by a data dictionary. Keep original evidence coordinates and normalized design coordinates in separate columns. If lift-over is necessary, record the source build, target build, tool, version, and intervals that failed or mapped ambiguously. A clean coordinate handoff reduces downstream confusion when results are delivered as VCF files for population genomics.
Add Flanks for a Reason
Flanking sequence serves different purposes. Short flanks may cover splice junctions. Longer flanks may support promoter analysis, haplotype reconstruction, or bait placement when the core target is difficult. Amplicon assays may need primer-placement sequence outside the analyzed interval, while hybrid capture can recover some sequence beyond a baited region. The physical enrichment footprint and the bioinformatic region of interest are therefore related but not identical.
For every flank rule, state:
- The biological reason for including it.
- Whether it is part of the reported analysis region.
- Whether it is expected to meet the same coverage threshold as the core target.
- How overlaps between adjacent intervals will be merged.
Avoid a universal "plus 500 bp" rule across exons, GWAS loci, QTLs, and enhancers. The same numerical window solves different problems poorly.
Screen Difficult and Low-Information Regions Early
Figure 3. Early sequence-context review identifies targets that may need alternative designs or qualified reporting.
Some targets are difficult to enrich, align, or interpret. Screen for repeats, low complexity, segmental duplications, paralogs and pseudogenes, extreme GC content, reference gaps, and common variants under proposed primer or bait sequences. In non-model organisms, also examine assembly contiguity and the risk that a coordinate represents a collapsed repeat or misassembly.
Classify each interval as straightforward, designable with caution, or unsuitable for the intended method. Required difficult loci may need alternate baiting, long-range amplification, orthogonal testing, or broader sequencing. Optional difficult loci may be removed. This triage should occur before final target-size calculations because nominal bases and reliably analyzable bases are not the same.
Balance Breadth, Depth, and Cohort Size
Target definition affects the whole project budget. Adding regions increases total target size and can reduce depth per sample at a fixed sequencing allocation. Increasing depth without changing the budget can reduce the number of samples. A broader panel also creates more multiple-testing burden and more variants requiring annotation and QC.
Use a scenario table rather than one optimistic estimate:
| Scenario | Target breadth | Depth objective | Cohort implication | Best use |
| Focused | High-priority exons or variants | Higher depth and tighter QC | Supports more samples per run | Rare variants in strong candidate regions |
| Balanced | Genes plus selected locus/regulatory intervals | Moderate-to-high depth | Intermediate multiplexing | Mixed coding and locus follow-up |
| Broad | Many genes, large QTLs, or extensive noncoding targets | Depth pressure increases | Fewer samples or more sequencing | Hypothesis set cannot be narrowed further |
The best scenario depends on allele frequency and analysis. Rare-variant discovery usually benefits from reliable depth and a sufficiently large cohort. An excessively broad panel can weaken both. Review the planned scale using the guide to population genomics sample size and sequencing depth.
Choose Enrichment After Defining the Targets
Do not let a preferred laboratory method determine the biological region list prematurely. First define the required targets, then compare amplicon and hybrid-capture feasibility. Amplicons can be efficient for compact, stable regions but face multiplex interactions and primer-site variation. Hybrid capture generally handles larger and more dispersed targets with flexible tiling, although duplicates, off-target reads, and nonuniform capture still require planning. The hybrid capture versus amplicon comparison covers this method decision without replacing region design.
If the final target footprint becomes too large or evidence remains diffuse, compare targeted resequencing with broader approaches. The targeted resequencing versus WGS guide addresses that boundary. If the panel is already validated and the challenge is production scale, use the separate guide on scaling targeted resequencing to large cohorts.
Pilot, Revise, and Lock the Target Version
A representative pilot should measure the fraction of target bases reaching the required depth, depth uniformity, on-target rate, duplicate rate, missingness, genotype concordance, and performance in difficult intervals. Review results at the region level, not only as a panel-wide average. A few large, easy intervals can hide systematic failure in short exons or high-priority regulatory elements.
After the pilot:
- Rescue required underperforming regions where technically possible.
- Remove optional regions that consume reads without usable coverage.
- Rebalance baits or primers when the method supports it.
- Freeze a versioned BED file, annotation table, transcript policy, and analysis-region file.
- Record every change between pilot and production.
Version locking protects multi-batch comparability. If a later panel version changes coordinates, bridge samples help separate biological differences from assay changes.
What Targeted Resequencing Cannot Resolve Alone
Targeted data can test variation inside the selected intervals, but it cannot recover causal candidates that were excluded by the target definition. This matters when a GWAS lead marker tags a broader haplotype, when regulatory evidence comes from another tissue or species, or when a QTL interval remains too wide to tile economically. Treat external annotations as evidence for prioritization, not universal proof that the same regulatory relationship operates in the study population. Preserve the excluded-region list so that negative results are interpreted against what was actually observable.
If a pilot shows uneven coverage, separate sequence-context failures from sample-wide failures. Recurrent gaps at the same interval suggest redesign, an alternate enrichment chemistry, or a qualified limitation. Broad failure in particular sample types suggests input quality, library preparation, or handling. Confirm important variants with an orthogonal method when mapping ambiguity, allele balance, or downstream decisions make false calls consequential. Finally, distinguish statistical association from mechanism: enrichment can refine a locus and identify candidates, but independent replication and functional experiments may still be required to establish biological causality.
Frequently Asked Questions
Only when intronic variation, splice regions, regulatory elements, haplotypes, or primer placement justify them. Full introns can expand the panel substantially.
There is no universal width. Prefer credible sets and population-relevant LD boundaries. Use a documented fixed window only when better evidence is unavailable.
Yes, but broad intervals may dominate the panel and reduce depth or sample scale. Narrowing and prioritization should be attempted first.
No. Include regulatory elements supported in the relevant biological context and label uncertain links clearly.
Provide the coordinate table, assembly and annotation versions, transcript policy, evidence and priority fields, variant classes, cohort plan, depth goal, and intended analyses.
References
- Mamanova L, Coffey AJ, Scott CE, et al. Target-enrichment strategies for next-generation sequencing. Nature Methods. 2010;7(2):111-118. doi:10.1038/nmeth.1419.
- Natsoulis G, Bell JM, Xu H, et al. A flexible approach for highly multiplexed candidate gene targeted resequencing. PLOS ONE. 2011;6(6):e21088. doi:10.1371/journal.pone.0021088.
- Lin H, Wang M, Brody JA, et al. Strategies to Design and Analyze Targeted Sequencing Data. Circulation: Cardiovascular Genetics. 2014;7(3):335-343. doi:10.1161/CIRCGENETICS.113.000350.
- Service SK, Teslovich TM, Fuchsberger C, et al. Re-sequencing expands our understanding of the phenotypic impact of variants at GWAS loci. PLOS Genetics. 2014;10(1):e1004147. doi:10.1371/journal.pgen.1004147.
- Schaid DJ, Chen W, Larson NB. From genome-wide associations to candidate causal variants by statistical fine-mapping. Nature Reviews Genetics. 2018;19(8):491-504. doi:10.1038/s41576-018-0016-z.
- Patel AP, Peloso GM, Pirruccello JP, et al. Targeted exonic sequencing of GWAS loci in the high extremes of the plasma lipids distribution. Atherosclerosis. 2016;250:63-68. doi:10.1016/j.atherosclerosis.2016.04.011.
- Bainbridge MN, Wang M, Wu Y, et al. Targeted enrichment beyond the consensus coding DNA sequence exome reveals exons with higher variant densities. Genome Biology. 2011;12(7):R68. doi:10.1186/gb-2011-12-7-r68.
For research purposes only. The information and services described here are not intended for clinical diagnosis, therapeutic decisions, or personal health assessment.