TCR-Guided Neoantigen Prioritization: When Repertoire Evidence Adds Value Beyond WES and RNA-seq
Figure 1. TCR evidence is an additional decision layer that can refine, but does not replace, tumor genomic and transcriptomic evidence.
Tumor whole-exome sequencing and RNA sequencing can identify expressed somatic alterations and generate HLA-matched peptide candidates. They do not directly show whether a candidate is processed, presented, recognized by a T cell, or associated with an expanding clonotype. TCR repertoire sequencing, antigen-specific receptor discovery, paired-chain measurements, and longitudinal tracking address different parts of that gap. Their value depends on study design: bulk repertoire data can reveal clonal dynamics but usually cannot assign antigen specificity by sequence alone, while direct receptor–peptide evidence is much stronger but more demanding to obtain. This guide explains how to use TCR evidence as a calibrated prioritization layer rather than a shortcut.
CD Genomics provides sequencing and bioinformatics for research use. The workflows described here are not clinical diagnostic or treatment-selection services.
Key Takeaways
- Keep candidate generation and receptor evidence separate. WES, RNA-seq, and HLA modeling define candidate peptide space; TCR assays test immune association at different strengths.
- Clonality is not specificity. An abundant or expanding receptor may be tumor related, bystander, pathogen reactive, or technically overrepresented.
- Paired receptor chains matter. Alpha and beta chains jointly determine most conventional TCR specificity; unpaired bulk sequences limit receptor reconstruction.
- Time and tissue provide context. Enrichment in tumor, expansion after stimulation, or reproducible longitudinal change is more informative than one cross-sectional frequency.
- Use an evidence ladder. Repertoire association, peptide–HLA binding, receptor binding, and functional response should not be merged into one unsupported score.
Start With the Genomic Candidate Set
TCR-guided prioritization begins with a defensible candidate set. Tumor-normal Whole Exome Sequencing supports somatic coding variant discovery. Tumor Transcriptome Sequencing tests gene and allele expression and can add fusion or splice-derived events. HLA Typing by Sequencing or another well-documented HLA source defines the peptide-presentation context.
Every candidate should retain its source variant or transcript event, mutant and wild-type peptide, HLA allele, prediction outputs, DNA support, RNA support, expression, and quality flags. HLA loss, tumor purity, clonality, transcript choice, and variant phasing may change interpretation. The TMB and HLA diversity guide explains why mutation quantity and presentation context are related but distinct variables.
Do not discard lower-ranked computational candidates before considering how the downstream assay works. Prediction algorithms can prioritize likely binders, yet immunogenicity also depends on antigen processing, peptide abundance, T-cell availability, tolerance, and receptor recognition. Systematic neoepitope–HLA discovery has shown that experimentally supported pairs can be organized across tumors, while also emphasizing the need for direct evidence rather than binding rank alone [1].
| Candidate layer | Question answered | Evidence produced | Question not answered |
|---|---|---|---|
| Tumor-normal WES | Is an alteration somatic and coding? | Variant reads and annotation | Is the mutant transcript expressed? |
| Tumor RNA-seq | Is the gene, allele, fusion, or junction observed? | Expression and transcript reads | Is the peptide presented? |
| HLA modeling | Which mutant peptides may bind submitted alleles? | Predicted rank or affinity | Does a TCR recognize the complex? |
| Immunopeptidomics | Is a peptide detected in an HLA-associated pool? | Spectrum-supported peptide evidence | Which receptor responds? |
| TCR repertoire | Which clonotypes are present or changing? | Frequency, diversity, and tracking | What antigen drives each clonotype? |
| Specificity or functional assay | Does a receptor associate with or respond to a defined target? | Pair-specific or response evidence | Is the relationship relevant in every tissue context? |
Decide Which TCR Evidence Is Needed
"Add TCR sequencing" can mean several different experiments. Bulk TCR sequencing measures receptor sequences and frequencies in a sample. Single-cell immune profiling can recover paired chains and transcriptional state. Antigen-stimulation or multimer-based enrichment can connect selected cells with a peptide–HLA target. Receptor reconstruction and functional assays can test a defined receptor more directly.
The appropriate method follows the decision the project needs to make. If the question is whether a clonotype expands after exposure to a candidate pool, longitudinal bulk data may be enough. If the question is which receptor recognizes one peptide–HLA complex, paired-chain and antigen-linked evidence are needed.
The BCR and TCR Sequencing service supports repertoire-focused measurements, and the single-cell RNA sequencing service can add cellular state and paired-chain context when the study design requires it. These assays should be scoped around a prespecified hypothesis rather than added after sequencing with no matched samples.
Figure 2. TCR evidence becomes more specific as the study progresses from repertoire association to receptor–target and functional measurements.
Understand What Bulk Repertoire Data Can Add
Bulk repertoire sequencing is well suited to measuring clonotype abundance, richness, diversity, sharing, and change over time. Tumor enrichment relative to blood, expansion after peptide stimulation, or consistent tracking across serial samples can identify clonotypes worth follow-up. The TCR sequencing workflow and applications guide describes the main laboratory and analysis stages.
Bulk evidence is most useful when samples are paired by participant and time point, processed with one protocol, and supported by enough template molecules. Molecular barcodes can reduce amplification bias in compatible workflows. Technical replicates or input-molecule estimates help distinguish a rare biological signal from sampling noise.
Useful comparisons include:
- tumor versus matched peripheral blood at the same time point;
- antigen-stimulated versus unstimulated aliquots from the same specimen;
- pre-exposure versus post-exposure time points with stable sampling conditions;
- candidate-positive versus candidate-negative tumor regions;
- sorted activated or multimer-positive cells versus their parent population;
- replicate libraries that test whether low-frequency clonotypes are reproducible.
The main caution is specificity. A clonotype can be abundant because of tissue residency, infection history, nonspecific activation, homeostatic expansion, or assay bias. Sequence similarity to a public receptor is not proof of identical antigen recognition. The immune repertoire sequencing methodologies guide can help align claims with assay resolution.
Recover Paired Chains and Cell State When Needed
Conventional alpha-beta TCR specificity is determined by the paired receptor, not the beta chain alone. Bulk assays often sequence loci separately and cannot determine which alpha chain belongs with which beta chain. This limitation matters when a receptor will be reconstructed, compared across cell states, or tested against a candidate.
Single-cell V(D)J plus transcriptome profiling can link paired receptor chains to cell phenotype, activation state, and tissue compartment. It can distinguish expanded cytotoxic-like, exhausted-like, memory-like, or other transcriptional states, while avoiding the assumption that all cells sharing a beta sequence are equivalent. Cell-state labels remain dependent on sampling, marker definitions, and analysis thresholds.
Design details include cell viability, recovery targets, doublet control, sequencing depth, sample multiplexing, V(D)J capture chemistry, and whether surface proteins or peptide–HLA reagents are added. A large cell count is not useful if the relevant population is lost during dissociation or enrichment. The immune repertoire multi-omics guide discusses how receptor sequence, expression, and other modalities can be combined.
Paired-chain evidence still does not establish antigen identity. It makes the receptor reconstructable and associates it with a cell, which is a necessary step toward stronger specificity testing.
Link Receptors to Candidate Antigens
Several experimental strategies can connect candidate peptides with T cells. Peptide pools followed by activation-marker sorting can identify responding populations but may initially localize specificity only to a pool. Peptide–HLA multimers can enrich binding cells when the correct HLA and reagent are available, although binding conditions and avidity influence detection. Genetic screens can test targets without relying exclusively on peptide–HLA binding prediction.
HLA-unbiased genetic screens have identified patient-specific CD4 and CD8 neoantigens [3]. Research combining neoantigen discovery with non-viral precision receptor replacement has further shown how candidate identification, TCR isolation, receptor reconstruction, and functional assessment can be connected [2]. These are multi-stage evidence programs, not outcomes produced by repertoire sequencing alone.
When candidate numbers are large, use staged testing:
- retain candidates that pass genomic identity, expression, and annotation review;
- group peptides into traceable pools with controls and balanced composition;
- screen for reproducible activation or enrichment relative to controls;
- deconvolute positive pools to individual peptide–HLA targets;
- recover paired receptor chains from responding cells;
- reconstruct or otherwise test the receptor–target relationship;
- retain negative and ambiguous results in the audit trail.
Sensitive discovery studies in solid tumors have demonstrated the value of integrating tumor antigen identification with cognate receptor recovery [4]. Assay controls should include irrelevant peptides, wild-type counterparts when appropriate, positive-response controls, HLA-mismatched or blocking conditions, replicate wells, and viability or background measurements.
Figure 3. A staged workflow preserves traceability from each genomic candidate to responding cells, paired receptors, and target-specific tests.
Use Computational TCR Models as Hypothesis Tools
TCR–antigen prediction models can help organize large search spaces, compare sequence motifs, or rank follow-up hypotheses. Pan-peptide meta-learning illustrates progress toward predicting receptor–antigen binding across peptides [5]. Model performance, however, depends on training-set composition, negative-example construction, HLA context, receptor-chain availability, and similarity between the query and known examples.
A high model score should not be treated as equivalent to experimental recognition. Public databases overrepresent common pathogens, selected HLA alleles, and receptors from particular assay systems. A receptor similar to a known sequence may share specificity, show cross-reactivity, or recognize something different.
Document the model name and version, input chains, peptide and HLA fields, training-data overlap checks, score calibration, and threshold. If a model uses only the beta CDR3 and peptide, its output should be labeled as sequence-based prioritization rather than complete receptor–peptide–HLA evidence.
Track Clonotypes Across Time and Compartments
Longitudinal evidence can add value when the project asks whether candidate-associated clonotypes appear, expand, contract, persist, or move between compartments. The clonotype-tracking guide explains why stable identifiers and comparable sampling are essential.
Define a clonotype consistently across time. Exact nucleotide, amino-acid CDR3, V/J gene, paired-chain, and clustering-based definitions are not interchangeable. A nucleotide-defined lineage can separate convergent rearrangements, while amino-acid aggregation may better reflect shared receptor sequence. Report both when the distinction affects interpretation.
Use absolute template estimates when possible and display sampling depth beside frequency. Apparent disappearance may mean the clone fell below detection because fewer cells or molecules were sampled. Batch, tissue processing, lymphocyte fraction, and repertoire input also affect observed dynamics.
Cross-compartment overlap is informative only when the biological and technical denominators are clear. A clonotype enriched in tumor and detectable in blood may be useful for tracking, but enrichment is not proof that it recognizes a tumor neoantigen. Antigen-linked or functional evidence is still required for that claim.
Design controls around the inference
Each proposed inference needs a control that can challenge it. For repertoire expansion, include an unstimulated or irrelevant-antigen condition and matched input. For peptide-specific enrichment, include wild-type peptides where biologically meaningful, HLA-mismatched reagents, and replicate staining or sorting. For receptor reconstruction, use negative receptors, positive assay controls, and target cells or antigen-presenting systems with verified HLA and antigen expression.
Background thresholds should be set without looking only at the favored candidate. Predefine the minimum cell count, replicate agreement, fold change, response magnitude, and viability required to call a screen positive. If many peptides or receptors are tested, account for the increased opportunity to observe an extreme value by chance. A result that appears in one of many wells and cannot be repeated belongs in an unresolved tier, not the highest evidence tier.
Negative findings also need power context. Failure to detect a clonotype can reflect inadequate cell sampling, low template count, chain dropout, peptide-presentation conditions, or a response below the assay limit. Record the number of input cells, recovered receptor molecules, usable paired cells, assay replicate counts, and control performance. This makes "not detected" distinguishable from "tested under conditions that could exclude the hypothesis."
Separate discovery and confirmation sets
When enough material is available, use one aliquot or time point to discover candidate-associated receptors and an independent aliquot, replicate, or assay to confirm them. Reusing the same measurements for nomination and validation inflates apparent confidence. A confirmation set should preserve the receptor and peptide definitions selected during discovery and apply prespecified thresholds without retuning them to the new result.
For scarce samples, an orthogonal confirmation method can provide a similar safeguard. For example, repertoire enrichment may nominate clonotypes, paired single-cell data may recover receptors, and a separate target-linked assay may test specificity. The evidence remains strongest when the analytical steps fail in different ways rather than repeating the same bias.
Build a Transparent Prioritization Matrix
The final rank should keep evidence dimensions visible. Avoid compressing all data into a single score without showing its components and missingness. A candidate with strong expression but no receptor evidence is different from one with moderate expression and a functionally tested paired receptor.
| Evidence dimension | Example measurement | Interpretation |
|---|---|---|
| Genomic confidence | Tumor-normal reads, allele fraction, caller agreement | Confidence that the alteration is real and somatic |
| Transcript evidence | Gene expression, mutant reads, junction support | Evidence that a relevant transcript is present |
| Presentation context | HLA call quality, binding rank, HLA loss status | Plausibility of peptide presentation |
| Repertoire association | Tumor enrichment, expansion, persistence | Association of clonotypes with context or time |
| Receptor specificity | Multimer, deconvolution, genetic-screen link | Evidence connecting receptor and candidate target |
| Functional response | Activation, cytokine, killing, or other prespecified readout | Assay-dependent evidence of response |
| Reproducibility | Replicates, orthogonal method, independent time point | Stability of the observed relationship |
| Uncertainty | Low input, missing chain, ambiguous HLA, sparse counts | Limits that should remain visible in ranking |
A simple tier system is often clearer than false precision. Tier 1 might require verified somatic and RNA evidence plus reproducible receptor–target and functional support. Tier 2 might have antigen-linked enrichment without complete functional confirmation. Tier 3 might contain strong genomic candidates with repertoire association only. The exact thresholds should be prospectively defined for the research program.
The ESCAPE-seq neoantigen research overview provides context for experimental screening after computational nomination. CD Genomics' immuno-oncology research solutions can integrate sequencing layers, but conclusions should stay within the evidence actually generated.
Figure 4. A transparent matrix preserves the difference between predicted, associated, antigen-linked, and functionally supported candidates.
Common Failure Modes
- Equating expansion with specificity: require antigen-linked evidence before assigning a target to an expanded clonotype.
- Using unpaired beta chains as complete receptors: state the limitation and recover paired chains before receptor reconstruction.
- Comparing unmatched tissues or batches: balance collection, processing, input, and sequencing so technical differences do not dominate.
- Ignoring negative results: retain candidates that were tested, assay conditions, and reasons a result was inconclusive.
- Overtraining on public specificity databases: check sequence and peptide similarity to training data and validate out-of-distribution predictions.
- Collapsing evidence into one rank: report each component, its threshold, and missing values.
FAQ
- Does an expanded TCR prove neoantigen recognition?
- Is bulk TCR sequencing useful without single-cell data?
- When does single-cell TCR sequencing add the most value?
- Can a prediction model replace functional validation?
- What samples should be collected prospectively?
References
- Gurung HR, Heidersbach AJ, Darwish M, Chan PPF, Li J, Beresini M, Zill OA, Wallace A, Tong AJ, Hascall D, Torres E, Chang A, Lou KHW, Abdolazimi Y, Hammer C, Xavier-Magalhães A, Marcu A, Vaidya S, Le DD, Akhmetzyanova I, Oh SA, Moore AJ, Uche UN, Laur MB, Notturno RJ, Ebert PJR, Blanchette C, Haley B, Rose CM. Systematic discovery of neoepitope–HLA pairs for neoantigens shared among patients and tumor types. Nature Biotechnology. 2024;42(7):1107-1117. doi:10.1038/s41587-023-01945-y
- Foy SP, Jacoby K, Bota DA, Hunter T, Pan Z, Stawiski E, Ma Y, Lu W, Peng S, Wang CL, et al. Non-viral precision T cell receptor replacement for personalized cell therapy. Nature. 2023;615(7953):687-696. doi:10.1038/s41586-022-05531-1
- Cattaneo CM, Battaglia T, Urbanus J, Moravec Z, Voogd R, de Groot R, Hartemink KJ, Haanen JBAG, Voest EE, Schumacher TN, Scheper W. Identification of patient-specific CD4+ and CD8+ T cell neoantigens through HLA-unbiased genetic screens. Nature Biotechnology. 2023;41(6):783-787. doi:10.1038/s41587-022-01547-0
- Arnaud M, Chiffelle J, Genolet R, Navarro Rodrigo B, Perez MAS, Huber F, Magnin M, Nguyen-Ngoc T, Guillaume P, Baumgaertner P, Chong C, Stevenson BJ, Gfeller D, Irving M, Speiser DE, Schmidt J, Zoete V, Kandalaft LE, Bassani-Sternberg M, Bobisse S, Coukos G, Harari A. Sensitive identification of neoantigens and cognate TCRs in human solid tumors. Nature Biotechnology. 2022;40(5):656-660. doi:10.1038/s41587-021-01072-6
- Gao Y, Gao Y, Fan Y, Zhu C, Wei Z, Zhou C, Chuai G, Chen Q, Zhang H, Liu Q. Pan-Peptide Meta Learning for T-cell receptor–antigen binding recognition. Nature Machine Intelligence. 2023;5(3):236-249. doi:10.1038/s42256-023-00619-3
- Kim JY, Cha H, Kim K, Sung C, An J, Bang H, Kim H, Yang JO, Chang S, Shin I, Noh SJ, Shin I, Cho DY, Lee SH, Choi JK. MHC II immunogenicity shapes the neoepitope landscape in human tumors. Nature Genetics. 2023;55(2):221-231. doi:10.1038/s41588-022-01273-y
For research use only. Not for use in diagnostic procedures or individual treatment decisions.