From BSA-Seq Interval to Candidate Gene
Your BSA-Seq, QTL-Seq, or MutMap analysis has done its job: it has converted a complex trait into one or more genomic intervals that track with the phenotype in your bulks. The hard part starts immediately after that.
A mapped peak is not a candidate gene conclusion. It's a search space that can still contain dozens—or hundreds—of genes and a long tail of variants that look "interesting" on paper.
This guide is intended for researchers who have identified candidate intervals through BSA-Seq, QTL-Seq, or MutMap and need a transparent, reproducible workflow for candidate gene prioritization. It explains how to convert a mapped interval into a ranked shortlist using defined evidence tiers and clear scoring criteria, with outputs that support downstream fine mapping and experimental validation.
Figure 1. Multiple evidence layers used to prioritize genes within a mapped interval.
Key Takeaway: Treat "the mapped interval" and "the prioritized candidate genes" as two different outputs. The first is positional. The second must be evidence-integrative and auditable.
What a BSA-Seq Interval Can and Cannot Tell You
Sequencing-based bulk segregant approaches (including QTL-Seq and MutMap) are efficient because they translate segregation into an allele-frequency signal along the genome. Recent reviews emphasize how these approaches accelerate locus discovery in plant systems, but also why locus discovery is not the same thing as gene discovery (see "Next-generation bulked segregant analysis for Breeding 4.0 (Cell Reports, 2023)" and "Harnessing the potential of bulk segregant analysis sequencing in plants (Frontiers in Genetics, 2022)").
A well-supported mapped interval provides a defensible region in which to concentrate downstream interpretation and validation. Its usefulness still depends on the quality of the original phenotype, population, sequencing, variant calling, and statistical analysis. It does not, by itself, establish which gene inside the region is causal.
Why Distance to the Peak Is Not Enough
Even when your peak looks sharp, the underlying signal is shaped by linkage, local recombination, and the specific variant density of the reference and population. In low-recombination regions, a long haplotype can carry many passenger variants; in regulatory architectures, the functional effect can act at a distance; and in many crop references, gene models are still incomplete enough that "nearest gene" is sometimes not the correct gene model.
A practical standard for this stage is simple: if your post-mapping logic cannot be written down as a set of rules that someone else can apply to the same interval and get the same ranked list, your result is not yet a usable candidate-gene deliverable.
BSA-Seq Candidate Gene Prioritization: Annotating Variants Within the Interval
The fastest way to shrink an interval gene list is to stop thinking in terms of "genes in the region" and start thinking in terms of gene–variant hypotheses. Candidate-gene work is about linking plausible causal variants to plausible causal genes.
In most projects, the first pass is small-variant annotation (SNPs and short indels) against a reference gene model. Tools differ, but the principle is the same: consequence annotation is most informative for coding/transcript effects and much weaker for non-coding inference.
For a concise, tool-agnostic framing of what consequence annotation means, the Ensembl VEP review "Annotating and prioritizing genomic variants using Ensembl VEP (Human Mutation, 2022)" is a clean citation. For a broader discussion of how variant-effect predictors behave across consequence classes, see "Variant effect predictors: a systematic review and practical guide (2024)".
SNP and indel functional annotation: what to record (not just what to compute)
Most "qtl interval gene annotation" summaries fail not because the tool is wrong, but because the output is not interpretable. To keep it auditable, record three things explicitly:
- Reference build and gene model version used for annotation.
- Rules for defining gene-associated non-coding regions (for example, what promoter window you used and whether you anchored it to TSS or transcript start).
- A gene-level synopsis, not only a variant list: "Which genes carry which consequence classes?"
This shifts your report from "here is a VCF" to "here is a prioritized hypothesis set."
High-Impact Coding Variants as Priority Sequence Evidence
Coding consequences that predict clear disruption of protein function, such as frameshift indels, stop-gained variants, and canonical splice donor or acceptor changes, are often assigned high priority when the calls are technically reliable and their segregation pattern is consistent with the mapped trait. They are not automatically causal, but they provide stronger sequence-layer evidence than consequence labels alone.
The failure mode is equally common: treating every missense variant as a top candidate. Missense changes can be causal, but they are abundant. Without additional evidence (conservation, domain disruption, haplotype consistency), missense-heavy genes can dominate your list by volume rather than by plausibility.
Promoter and putative regulatory variants (use conservative language)
Regulatory variants matter in crops, but "promoter variant" is often a location label rather than a functional claim. If you want a reproducible framework, split regulatory evidence into two classes:
- Putative regulatory: upstream/UTR/intronic variants flagged by position relative to gene models.
- Regulatory with support: variants that also align with independent evidence (trait-relevant expression shifts, motif disruption with plausible biology, or regulatory datasets).
This is the practical takeaway from modern discussions of variant predictors: outside coding regions, consequence labels are hypothesis-generating, not verdicts.
Structural variants (SVs): include them as a first-class evidence layer
In many crop systems, SVs—presence/absence variation, copy-number changes, inversions, large insertions/deletions—can produce large phenotypic effects while being poorly captured by SNP-only workflows. Even if you do not run a full SV discovery workflow at this stage, you should still treat SVs as a ranked evidence layer: "SV evidence present," "SV evidence absent," or "SV evidence not assessed."
⚠️ Warning: A candidate list that ignores SVs can be systematically biased toward genes that are easier to annotate rather than genes that are more likely causal.
If your project needs an end-to-end workflow that standardizes interval definition and variant annotation, the appropriate upstream route may include BSA mapping services, MutMap services, or a broader QTL mapping service.
Integrating Expression and Functional Evidence
Variant evidence is often necessary to shrink lists, but it is rarely sufficient to explain mechanism—especially in polygenic traits, regulatory architectures, or regions with dense polymorphism. This is where expression and functional interpretation become decisive, provided you treat them as support rather than as single-factor proof.
Figure 2. An integrated evidence framework for candidate gene prioritization.
RNA-Seq Differential Expression
Integrating RNA-seq with mapping is now common in crop pipelines and is often presented as a way to narrow candidate genes inside mapped intervals. A recent example of this combined pattern (QTL mapping + BSA-seq + RNA-seq) is described in "Integrating QTL mapping, BSA-seq and RNA-seq to identify candidate genes… (Frontiers in Plant Science, 2025)".
Use DE correctly by asking one practical question: does the expression signal strengthen the same causal story implied by segregation and sequence?
- If a gene is inside the interval, carries a plausible coding/regulatory variant, and is also differentially expressed in the trait-relevant tissue/time, the hypothesis becomes sharper.
- If a gene is differentially expressed but carries no plausible variant signal (and sits in a dense LD block), treat it as a downstream-response hypothesis unless other evidence elevates it.
When new expression data are required, Transcriptome RNA-Seq Services can generate trait-relevant expression profiles, while Agricultural Transcriptomic Data Analysis can support differential expression, pathway analysis, and integration with the mapped interval.
Tissue and Stage Specificity as a Plausibility Filter
A deceptively strong filter is simply verifying that a candidate gene is expressed where/when the trait is determined. This is rarely enough to promote a gene to "high priority," but it is often enough to demote a gene that is biologically implausible.
The key is to document the rule you used—TPM threshold, rank-based expression, or presence/absence—so this layer is reproducible.
Gene ontology and pathway context
GO and pathway evidence are valuable because they help translate a gene ID into a mechanistic story that your team can reason about. But they are also a source of overconfidence when used alone.
The most defensible use is: treat GO/pathway as contextual evidence that can strengthen a candidate when paired with segregation/variant signals, not as a primary driver that substitutes for genetic evidence.
Orthologs, domains, conservation, and literature
This is the "plausibility amplifier" layer. A strong ortholog story, a conserved domain hit, or relevant trait literature can substantially increase confidence in a candidate when it aligns with the genetic and variant signals.
A useful, modern citation for the general principle of combining evidence types in candidate regions is "A bioinformatics toolbox to prioritize causal genetic variants in candidate regions (Trends in Genetics, 2024)". While not crop-specific, its argument is directly relevant: candidate-region interpretation is strongest when multiple annotation and evidence domains converge.
Using Haplotypes and Comparative Genomics
For many crop projects, haplotype evidence is what turns a long interval list into a short action list.
Haplotype and allele-frequency evidence
At the interval stage, you already know allele frequencies diverge across bulks. At the candidate stage, you need to show which genes and variants carry the most coherent segregation story.
In practice, strong haplotype evidence has two features:
- Consistency: the allele or haplotype block is consistently enriched in the high bulk and depleted in the low bulk (and, ideally, replicated).
- Resolution: recombination or additional genotyping narrows the interval in a way that keeps the candidate gene inside the boundary.
This is where your downstream plan matters. If you are moving toward fine mapping, you often need additional markers and recombinant screening; understanding how recombination shapes mapping resolution is part of keeping this interpretable (see genetic linkage and recombination for a refresher on how these forces interact). When additional recombinant screening is required to refine the boundaries, follow the workflow described in Fine-Mapping After QTL-Seq.
Comparative genomics (supporting, not decisive)
Comparative evidence is useful when it clarifies function (for example, an ortholog in a related species is known to regulate a pathway tightly tied to your trait). It should not replace segregation or variant evidence.
In a ranking rubric, comparative genomics usually belongs as supporting or contextual evidence—strong only when it agrees with other layers.
Building a Candidate Gene Ranking Matrix
If you want a candidate gene list that survives handoffs (to a marker-development team, a fine-mapping effort, or a functional validation group), you need a transparent ranking artifact.
Evidence tiers
A practical four-tier system works well because it forces you to declare what counts as decisive:
- Strong evidence: genetic/segregation support plus a high-confidence functional variant signal (high-impact coding or compelling regulatory/SV evidence) that fits the phenotype story.
- Supporting evidence: one strong layer (e.g., DE in the right context) plus at least one additional plausibility layer (e.g., regulatory variant or conserved-domain disruption).
- Contextual evidence: functional plausibility signals (GO/pathway/ortholog/tissue expression) that do not yet align with strong variant/segregation evidence.
- Weak or nonspecific evidence: proximity-only, DE-only in irrelevant context, or broad functional annotations that do not narrow.
Ranking Table Example
The following hypothetical example illustrates how the matrix can be completed. It is not a universal scoring rule.
| Illustrative Gene | Variant Evidence | Expression Evidence | Functional Evidence | Haplotype Evidence | Priority |
|---|---|---|---|---|---|
| Gene A | Predicted frameshift variant with reliable read support | Differentially expressed in the trait-relevant tissue | Conserved domain and known pathway relevance | Favorable allele retained within the refined haplotype block | High |
| Gene B | Putative promoter variant | Tissue-specific expression but no significant differential expression | Ortholog associated with a related process | Partial support across the candidate interval | Medium |
| Gene C | No compelling coding, regulatory, or SV evidence | Differential expression only in a non-relevant condition | Broad GO annotation | Outside the most strongly supported haplotype segment | Low |
Figure 3. Example evidence-based ranking of candidate genes.
How to Keep the Matrix Auditable
Write each cell so a reviewer can understand what you did without opening your full pipeline:
- Variant Evidence: list the top candidate variant(s) and consequence class (and whether SV evidence exists).
- Expression Evidence: specify tissue/time/condition and direction; or state "expressed in relevant tissue; not DE in tested condition."
- Functional Evidence: domain notes, pathway alignment, and the single strongest literature hook.
- Haplotype Evidence: allele-frequency directionality and any recombinant boundary support.
Why concordance is more reliable than any single proxy
Single signals are easy to misread: DE can be downstream, annotation can be noisy in poor gene models, and distance-to-peak is often a reflection of local LD rather than biology. When multiple independent evidence layers point to the same gene, you have something closer to a causal model rather than a convenient guess.
This is the core logic behind modern prioritization thinking: the goal is not to "find a gene," but to produce the smallest set of genes whose combined evidence justifies expensive downstream tests.
Common Prioritization Mistakes
The most common mistakes are predictable:
- Treating the mapped interval as the final biological conclusion.
- Using "closest gene to peak" as the deciding rule.
- Treating differential expression as proof of causality.
- Treating functional annotation as proof of causality.
- Ignoring SVs because they are harder to call.
- Failing to document promoter window definitions, gene model versions, or ranking rules.
If you correct only one habit, correct this: distance, annotation, and DE are not proofs. They are evidence layers that need to agree with segregation and variant logic.
Recommended Deliverables
A good candidate-gene package should be designed for downstream execution, not just for reporting. At minimum, it should include:
- interval coordinates with reference build and gene model version
- an annotated variant set for the interval, plus a short gene-level summary of consequence classes
- the ranking matrix (gene-by-evidence table) and the evidence-tier rubric you used
- a short, decision-focused narrative on the top-ranked candidates (typically top 3–10)
If your next step is marker development or validation genotyping, it helps to align the ranked candidates with the downstream marker strategy (see GBS-based marker-assisted selection). If the population design is a bottleneck for resolution, it may also be useful to revisit which populations best support your mapping and validation goals (see common genetic and breeding populations). For projects building integrated maps to support fine mapping, a dedicated genetic linkage map service can be part of that workflow.
FAQ
Is the gene closest to the BSA-Seq peak usually the causal gene?
No. A peak reflects linkage and allele-frequency divergence across the region, not a direct measure of functional causality. In low-recombination regions and dense haplotype blocks, the maximum statistic can be offset from the true functional variant and its target gene. Distance can be used as a weak tie-breaker inside a very narrow interval, but it should never be your sole selection rule.
Can RNA-seq differential expression prove a candidate gene is causal?
No. Differential expression is useful supporting evidence, but it is conditional on tissue, time point, and environment, and it can capture downstream responses rather than the causal mechanism. A causal gene may also show no detectable differential expression under the sampled conditions. Treat RNA-seq as a way to strengthen a hypothesis when it agrees with segregation and variant evidence.
If a gene has a predicted high-impact variant, is it automatically top priority?
Not automatically. High-impact coding variants are strong sequence-layer signals, but they can still arise from annotation artifacts, background mutations, or isoform-specific effects that do not match the trait biology. The most defensible prioritization occurs when high-impact variants also show allele-frequency or haplotype patterns consistent with the trait and when expression/function context does not contradict the hypothesis.
What does qtl candidate gene identification look like when the interval is still large?
It usually becomes a structured triage problem. You annotate all variants in the interval, summarize evidence at the gene level, then rank genes using a fixed rubric that weights strong sequence/haplotype evidence highest and uses expression and functional interpretation as supporting layers. The output is a shortlist that is small enough to carry into fine mapping, marker validation, or functional testing without reinterpreting the entire interval from scratch.
What should I do after I have a ranked candidate list?
Use the ranking to plan the next narrowing step rather than jumping straight into expensive biology. Common next steps are: refining boundaries with recombinants and additional markers, validating top variants across individuals (not only bulks), and selecting a minimal set of genes for functional tests. The goal is to convert a ranked list into one or two causal models with clear predictions.
Conclusion
A mapped interval from BSA-Seq, QTL-Seq, or MutMap is a high-value output, but it is not a causal answer. A credible candidate gene prioritization framework makes your post-mapping logic explicit: it integrates variant consequence (including SVs), expression context, haplotype consistency, and functional plausibility—and it promotes candidates based on concordance rather than convenience.
In other words, crop candidate gene analysis is a documentation problem as much as a biology problem: you're building a shortlist that others can reproduce, critique, and validate.
Discuss integrated candidate-gene prioritization for your mapped interval. Learn more about the QTL mapping service.
For research purposes only, not intended for clinical diagnosis, treatment, or individual health assessments.
Send a MessageFor any general inquiries, please fill out the form below.






