Three Neoantigen Discovery Strategies for Personalized mRNA Cancer Vaccine Research
Figure 1. Three complementary evidence layers support neoantigen discovery from coding mutations, broader genomic events, and noncanonical translation.
Personalized mRNA cancer vaccine research begins with a difficult selection problem: which tumor-specific sequences are sufficiently supported to become meaningful neoantigen candidates? A mutation list alone is not enough. Candidate antigens may arise from coding SNVs and InDels, gene fusions and structural events, abnormal transcript isoforms, or translated regions outside conventional protein-coding annotations.
The challenge is therefore not simply to identify as many alterations as possible. It is to build an evidence chain that progressively connects a tumor-specific event to a biologically plausible antigen source. Depending on the project, that chain may begin with coding mutations, expand to genome-wide structural events, or extend further into full-length transcript structures and active translation.
This article examines three complementary neoantigen discovery strategies and the published evidence behind them: WES plus RNA-seq for conventional mutation-derived candidates, WGS plus WES plus RNA-seq for broader genomic event discovery, and WES plus long-read RNA-seq plus Ribo-seq for transcript- and translation-aware discovery.
For an integrated research workflow that continues from neoantigen discovery into sequence design, plasmid preparation, mRNA production, and mRNA-LNP formulation, see the Personalized mRNA Cancer Vaccine Development Service.
TL;DR
- WES plus RNA-seq is a practical starting point for expressed coding SNVs, InDels, and selected fusion candidates.
- Adding WGS broadens discovery toward structural variants, rearrangements, fusion breakpoints, abnormal junctions, and selected noncoding events.
- Long-read RNA-seq improves full-length transcript resolution, while Ribo-seq adds translation-level evidence for noncanonical and cryptic antigen sources.
- Different technologies answer different questions: DNA defines the alteration, RNA indicates transcript support, long reads clarify transcript structure, and Ribo-seq indicates translation.
- No sequencing layer alone proves peptide presentation or T-cell recognition; each method contributes a different part of the evidence chain.
Why Neoantigen Discovery Requires Multiple Evidence Layers
A genomic alteration is only the first step toward a potential neoantigen. The altered sequence must exist in the tumor, be transcribed or otherwise generate an antigenic source, and ultimately produce peptides that can enter downstream presentation and immune-recognition pathways.
A useful evidence chain is: DNA alteration → RNA expression → transcript structure → translation → peptide presentation → T-cell recognition.
WES focuses mainly on coding genomic regions and is well suited to somatic SNVs and small InDels. Tumor-normal comparison helps separate somatic events from inherited background, while tumor RNA-seq adds expression support and can reveal selected fusion or splice-junction events. This combination is therefore well aligned with projects focused on conventional mutation-derived neoantigens.
WGS broadens genomic discovery to structural variants, rearrangements, noncoding regions, copy-number changes, and fusion breakpoints. This matters because some potentially useful antigen sources arise from altered genomic architecture rather than a simple nucleotide substitution within an annotated coding exon.
Long-read RNA sequencing adds a different type of information. Instead of reconstructing transcript structure from many short fragments, full-length or near-full-length reads can preserve exon connectivity across abnormal isoforms. This is particularly useful when multiple splice junctions, transcript fusions, alternative start sites, or noncanonical transcript structures must be assigned to the same RNA molecule.
Ribo-seq then moves the evidence from transcript presence toward active translation. Ribosome-protected fragments can reveal translated open reading frames that may not appear in standard protein annotations. This becomes especially relevant when conventional coding mutations yield few candidates or when the research question includes cryptic and noncanonical antigen sources.
These layers should be viewed as complementary rather than interchangeable. More sequencing is not automatically better; the research question should determine which evidence layers are needed and how much uncertainty each layer is expected to reduce.
Figure 2. Neoantigen discovery can progress from DNA alterations to transcript evidence and translation-level support before downstream presentation and immune-recognition studies.
What Each Evidence Layer Contributes to Candidate Confidence
One useful way to think about neoantigen discovery is as progressive candidate filtering. Each evidence layer removes a different type of uncertainty.
| Evidence layer | Main question answered | Typical contribution | What remains uncertain |
|---|---|---|---|
| Matched tumor-normal DNA | Is the alteration tumor-associated rather than inherited? | Defines somatic SNVs, InDels, and broader genomic events depending on the assay | Whether the event is transcribed or translated |
| Tumor RNA-seq | Is the altered event represented in the tumor transcriptome? | Adds expression, fusion, and selected junction evidence | Full transcript structure and active translation |
| Long-read RNA-seq | What is the complete structure of the abnormal transcript? | Resolves full-length isoforms, exon connectivity, junctions, and transcript fusions | Whether the transcript produces a translated antigen source |
| Ribo-seq | Is the region engaged by translating ribosomes? | Adds translation-level support and reveals noncanonical ORFs | Whether the resulting peptide is presented and immunogenic |
| Downstream presentation or immune-recognition studies | Is the peptide presented and biologically recognized? | Adds orthogonal evidence beyond sequencing | Context-dependent biological relevance |
This framework explains why a candidate supported by DNA, RNA, and translation evidence is not equivalent to a candidate supported by DNA alone. It also explains why projects should not automatically discard conventional WES plus RNA-seq when coding mutations are already the dominant biological question.
Case Study 1: WES + RNA-seq for Mutation-Derived Neoantigen Discovery
Research question
Can expressed, patient-specific tumor mutations be converted into individualized RNA vaccine candidates capable of stimulating measurable neoantigen-specific T-cell responses?
Published study
Rojas LA, Sethna Z, Soares KC, et al. Personalized RNA neoantigen vaccines stimulate T cells in pancreatic cancer. Nature. 2023;618:144-150. View the published study.
Study design
The study used surgically resected pancreatic ductal adenocarcinoma material to build individualized neoantigen vaccine designs. Patient-specific tumor and normal material supported somatic mutation identification, while tumor RNA information added expression evidence for candidate selection. The study then incorporated prioritized candidates into individualized RNA vaccine constructs.
The discovery logic is especially relevant to a classic neoantigen workflow because it demonstrates how genomic and transcriptomic evidence can be used sequentially rather than independently. Tumor-normal DNA comparison narrows the search to somatic alterations. RNA information then helps focus attention on altered sequences that are represented in the tumor transcriptome before construct design begins.
In the published trial, the individualized vaccine could contain up to 20 neoantigens per patient. This illustrates an important practical point for discovery: the upstream sequencing workflow often produces more possible candidates than can be taken forward, so prioritization is an essential part of the process rather than an optional downstream step.
Figure 3. WES plus tumor RNA-seq combines somatic coding alterations with expression evidence before candidate neoantigen prioritization.
Key findings
Sixteen participants received the individualized vaccine within the reported study workflow. Eight of the sixteen developed de novo, high-magnitude neoantigen-specific T-cell responses. Across patients evaluable at the single-target level, 25 of 230 administered vaccine neoantigens produced responses detectable by the study's high-threshold ex vivo assay, illustrating that only a subset of computationally selected candidates became strongly immunogenic in that experimental context.
Half of the responders recognized more than one vaccine neoantigen. The investigators also observed vaccine-associated T-cell clonal expansion and persistence, providing orthogonal evidence that the selected neoantigens were biologically relevant in a subset of participants.
These are findings from an independent published study and should not be interpreted as service performance data. The value of the study here is methodological: it connects tumor sequencing, candidate selection, individualized RNA construct design, and downstream immune-response measurement in a single research framework.
What this study demonstrates
For projects centered on conventional mutation-derived neoantigens, WES plus tumor RNA-seq provides a practical discovery foundation. WES identifies coding somatic alterations, while RNA sequencing adds evidence that selected events are expressed. This combination can reduce a large mutation list into a more biologically relevant candidate pool.
The study also highlights why candidate ranking matters. Even after genomic and transcriptomic filtering, not every selected peptide becomes strongly immunogenic. Upstream sequencing therefore improves candidate quality, but it does not eliminate the need for downstream prioritization and biological validation.
The remaining limitation is equally important: RNA expression does not directly establish translation or peptide presentation, and an exome-centered strategy is less complete for large structural, complex rearrangement, noncoding, or noncanonical antigen sources.
Case Study 2: Expanding Neoantigen Discovery Beyond Coding Mutations
SNVs and small InDels are not the only potential sources of tumor neoantigens. Structural changes can create novel junction sequences, fusion proteins, altered reading frames, and expression patterns that do not exist in the corresponding normal genome. For tumors where such events are important, broader genomic profiling can substantially expand the candidate space.
Research evidence 1: Gene fusion-derived neoantigens
Yang W, Lee KW, Srivastava RM, et al. Immunogenic neoantigens derived from gene fusions stimulate T cell responses. Nature Medicine. 2019;25:767-775. View the published study.
The researchers investigated tumors in which conventional mutation burden alone did not fully explain immune reactivity. Using genome-wide and RNA sequencing, they identified fusion events that created novel junction-derived peptide sequences. In an exceptional responder with metastatic head and neck cancer, a DEK-AFF2 fusion generated a peptide capable of eliciting a cytotoxic T-cell response.
The study also evaluated additional fusion-positive tumors and provided evidence that gene fusion-derived neoantigens can be subject to immune surveillance. The broader lesson is that a tumor with relatively few conventional coding mutations may still contain biologically important antigen sources created by structural rearrangement.
For neoantigen discovery, fusion candidates are especially interesting because the breakpoint can generate sequence combinations absent from normal proteins. Detecting the genomic event and confirming the corresponding transcript therefore provides a stronger basis for candidate reconstruction than either DNA or RNA alone.
Research evidence 2: The value of WGS + WES + RNA-seq
Newman S, Nakitandwe J, Kesserwan CA, et al. Genomes for Kids: The Scope of Pathogenic Mutations in Pediatric Cancer Revealed by Comprehensive DNA and RNA Sequencing. Cancer Discovery. 2021;11:3008-3027. View the published study.
The prospective study evaluated 309 children using WGS, WES, and RNA sequencing. It was not a neoantigen vaccine trial, but it provides a clear demonstration of what WGS contributes when added to exome and transcriptome profiling.
Across the cohort, 86% of patients harbored variants with diagnostic, prognostic, therapeutic, or cancer-predisposition relevance. Importantly for method selection, inclusion of WGS enabled detection of activating gene fusions in 36% of tumors, enhancer hijacking events in 8%, and small intragenic deletions in 15%, in addition to other variant classes.
These percentages should not be interpreted as expected neoantigen yields. Their value is to illustrate the range of genomic events that can emerge when the search expands beyond coding exons. In a neoantigen context, some of those event classes may generate novel junctions, altered reading frames, or tumor-associated transcripts that warrant further investigation.
Figure 4. WGS, WES, and RNA-seq broaden candidate discovery toward coding mutations, fusions, structural events, rearrangements, and transcript-supported junctions.
What these studies demonstrate
The main value of a WGS plus WES plus RNA-seq strategy is expanded event discovery rather than sequencing depth alone. Genome-wide profiling can broaden candidate exploration toward structural variants, rearrangements, fusion breakpoints, abnormal junctions, copy-number-associated events, and selected noncoding alterations.
WES remains useful within this combined strategy because it provides focused coverage of coding regions, while RNA evidence can determine whether selected genomic events are reflected at the transcript level. The three data layers therefore contribute different but complementary evidence.
The trade-off is a larger interpretation burden. More candidate sources require more filtering, annotation, prioritization, and downstream validation. A genome-wide strategy is most valuable when the tumor biology or research hypothesis justifies that broader search space.
Case Study 3: Translation-Aware Discovery of Noncanonical and Cryptic Antigens
The third strategy addresses a different limitation. A candidate antigen does not need to originate from a conventional annotated protein-coding mutation. Cancer cells can generate abnormal transcript isoforms and translated open reading frames outside standard annotations, creating antigen sources that are poorly represented by conventional mutation-centric pipelines.
Evidence 1: Long-read RNA sequencing resolves abnormal transcript sources
Oka M, Xu L, Suzuki T, et al. Aberrant splicing isoforms detected by full-length transcriptome sequencing as transcripts of potential neoantigens in non-small cell lung cancer. Genome Biology. 2021;22:9. View the published study.
Using full-length cDNA sequencing across 22 cancer cell lines, the researchers identified 2,021 novel splicing isoforms. Some isoforms showed protein-level support, and selected abnormal transcript structures produced predicted neoantigen candidates. The team also detected aberrant splicing isoforms in seven non-small-cell lung cancer specimens.
The study provides a direct example of why transcript structure matters. Short-read sequencing can identify individual splice junctions, but assigning distant junctions to the same full-length transcript can be ambiguous. Long-read sequencing can preserve exon connectivity and therefore improve reconstruction of the exact abnormal sequence that may produce a candidate antigen.
The authors also reported that approximately half of the peptide candidates evaluated in an ex vivo T-cell response assay showed the potential to activate T-cell responses through HLA interaction. This result does not mean that half of all long-read-derived isoforms are immunogenic; rather, it illustrates that selected transcript-derived candidates can progress beyond purely computational prediction.
Evidence 2: Ribo-seq identifies translated noncanonical ORFs
Ouspenskaia T, Law T, Clauser KR, et al. Unannotated proteins expand the MHC-I-restricted immunopeptidome in cancer. Nature Biotechnology. 2022;40:209-217. View the published study.
The researchers used ribosome profiling to identify translated novel and unannotated open reading frames and integrated those data with MHC-I immunopeptidomics. They reported 3,555 translated novel ORFs represented in the MHC-I immunopeptidome, demonstrating that antigen sources can extend well beyond conventional protein-coding annotations.
For neoantigen discovery, the important distinction is that Ribo-seq adds evidence of active translation. RNA abundance alone cannot show whether a transcript region is actually used by ribosomes. Translation-aware profiling therefore helps separate merely transcribed regions from regions with direct evidence of ribosome engagement under the tested conditions.
Evidence 3: Cryptic antigens can be recognized by T cells
Ely ZA, Kulstad ZJ, Gunaydin G, et al. Pancreatic cancer-restricted cryptic antigens are targets for T cell recognition. Science. 2025;388:eadk3487. View the published study.
The study extended the evidence chain from noncanonical translation toward presentation and immune recognition. Using high-resolution immunopeptidomics, the investigators found that cryptic peptides were abundant in the pancreatic cancer immunopeptidome and that approximately 30% of the examined noncanonical HLA-I-bound peptides showed cancer-restricted translation.
Selected cancer-restricted candidates demonstrated immunogenic potential in an ex vivo T-cell priming platform. The study also identified antigen-reactive TCR clonotypes and showed that selected redirected T cells could recognize patient-derived pancreatic cancer organoids expressing endogenous levels of the target cryptic antigens.
This is a useful reminder that translation is only one layer of evidence. The biological chain becomes progressively stronger as a candidate moves from abnormal transcript structure to active translation, peptide presentation, and finally immune recognition.
Figure 5. Long-read RNA-seq resolves abnormal transcript structures while Ribo-seq adds translation-level evidence for noncanonical and cryptic antigen candidates.
What these studies demonstrate
The WES plus long-read RNA-seq plus Ribo-seq strategy is designed to extend discovery beyond conventional coding alterations. WES still provides tumor-associated DNA variation, long-read RNA sequencing clarifies full-length transcript and junction structures, and Ribo-seq adds evidence that selected regions are actively translated.
Together, these layers can support candidate discovery from abnormal splice isoforms, noncanonical ORFs, cryptic translation events, and other transcript-derived sources. This is especially relevant when conventional coding candidates are limited or when the research question explicitly includes noncanonical antigen biology.
However, ribosome occupancy is not equivalent to peptide presentation, and peptide presentation does not automatically establish immunogenicity. The methods should therefore be interpreted as progressively stronger evidence layers rather than interchangeable proof.
Three Neoantigen Discovery Strategy Workflows
Classic strategy: WES + RNA-seq
The classic strategy compares tumor and matched-normal DNA to identify somatic coding alterations, then integrates tumor RNA evidence to prioritize expressed candidates. It is most direct when the expected antigen sources are coding SNVs, InDels, and selected fusion or transcript events.
Figure 6. Classic strategy: tumor and matched-normal WES identify somatic SNVs and InDels, while tumor RNA-seq adds expression and fusion evidence before candidate prioritization.
Whole-genome strategy: WGS + WES + RNA-seq
The whole-genome strategy combines genome-wide event discovery with focused coding-region analysis and tumor transcript evidence. WGS broadens structural and noncoding event detection, WES reinforces coding-region characterization, and RNA-seq helps determine whether selected events are represented at the transcript level.
Figure 7. Whole-genome strategy: WGS expands structural and rearrangement discovery, WES reinforces coding variants, and RNA-seq adds transcript-level evidence.
Integrated translatome strategy: WES + Long-Read RNA-seq + Ribo-seq
The integrated translatome strategy combines somatic DNA alterations with full-length transcript structures and ribosome-associated translation evidence. The purpose is not simply to add more assays, but to move candidate evaluation from DNA through transcript architecture to active translation.
Figure 8. Integrated translatome strategy: WES, long-read RNA-seq, and Ribo-seq contribute DNA, transcript-structure, and translation evidence to candidate neoantigen prioritization.
Comparing Three Neoantigen Discovery Strategies
The strategies should not be interpreted as simple tiers where more data are always better. Each combination addresses a different type of biological uncertainty and expands the candidate landscape in a different way.
| Strategy | Main evidence layer | Candidate sources | Primary strength | Remaining limitation |
|---|---|---|---|---|
| WES + RNA-seq | DNA + expression | Coding SNVs, InDels, selected fusions | Practical foundation for expressed mutation-derived candidates | Limited coverage of complex structural and noncoding events |
| WGS + WES + RNA-seq | Genome + exome + transcript | SNVs, InDels, fusions, SVs, rearrangements, junctions | Expands the genomic candidate search space | Greater interpretation and validation burden |
| WES + Long-Read RNA-seq + Ribo-seq | DNA + transcript structure + translation | Coding alterations, abnormal isoforms, noncanonical ORFs, cryptic translation | Adds transcript-resolution and translation evidence | Translation does not prove peptide presentation or T-cell recognition |
How to Choose a Neoantigen Discovery Strategy
When coding mutation-derived candidates are the main focus
WES plus RNA-seq is a practical starting point when the research question centers on expressed coding SNVs and InDels. The combination provides somatic mutation discovery together with tumor expression evidence and keeps the analysis focused on conventional mutation-derived candidates.
When fusions, structural variants, or complex events may matter
WGS plus WES plus RNA-seq is more appropriate when the expected antigen sources include rearrangements, fusion breakpoints, structural events, abnormal junctions, or broader genome-wide alterations. It is particularly useful when the tumor biology suggests that important events may occur outside conventional coding exons.
When transcript structure is itself part of the biological question
Long-read RNA sequencing becomes valuable when abnormal splicing, transcript fusions, alternative isoforms, or complex exon connectivity may influence the antigen sequence. It is not simply a deeper version of short-read RNA-seq; it contributes direct information about transcript architecture.
When conventional coding candidates are limited
Ribo-seq can add value when the research question extends toward noncanonical ORFs or cryptic translation. Translation evidence can help deprioritize transcript-derived candidates that lack detectable ribosome engagement under the tested conditions and can reveal translated regions outside standard annotations.
The goal is not to maximize candidate count. The goal is to generate a candidate set supported by evidence that matches the downstream research question.
Figure 9. A strategy-selection guide linking biological questions to coding, structural, transcript, and translation evidence.
How Discovery Strategy Influences Downstream mRNA Vaccine Research
Neoantigen discovery affects more than candidate ranking. The evidence supporting each candidate also influences how confidently researchers can trace the final construct back to the originating tumor event.
A coding SNV candidate may be supported by matched tumor-normal DNA and tumor RNA expression. A fusion-derived candidate may require breakpoint-level genomic evidence plus transcript confirmation. A noncanonical candidate may depend on full-length transcript reconstruction and translation-level support. These are different evidence histories even if all three candidates eventually enter the same downstream sequence-design workflow.
This traceability becomes especially important when multiple candidate classes are combined in one research construct. Recording the source event, transcript context, translation evidence, and candidate-prioritization rationale helps preserve the connection between the tumor biology and the final mRNA design.
Researchers requiring an integrated workflow can continue from multi-omics neoantigen discovery into sequence design, plasmid preparation, research mRNA production, and mRNA-LNP formulation through the Personalized mRNA Cancer Vaccine Development Service.
FAQ
- What is the difference between neoantigen discovery and neoantigen prediction?
- Is WES + RNA-seq sufficient for personalized neoantigen discovery?
- When should WGS be added?
- Does long-read RNA sequencing replace short-read RNA-seq?
- What does long-read RNA sequencing add to neoantigen discovery?
- Does Ribo-seq confirm that a candidate is a true neoantigen?
- Why is matched-normal sequencing important?
- Can coding, fusion-derived, and noncanonical candidates be evaluated in the same project?
References
- Lang F, Schrörs B, Löwer M, et al. Identification of neoantigens for individualized therapeutic cancer vaccines. Nature Reviews Drug Discovery. 2022;21:261-282. DOI.
- Rojas LA, Sethna Z, Soares KC, et al. Personalized RNA neoantigen vaccines stimulate T cells in pancreatic cancer. Nature. 2023;618:144-150. DOI.
- Yang W, Lee KW, Srivastava RM, et al. Immunogenic neoantigens derived from gene fusions stimulate T cell responses. Nature Medicine. 2019;25:767-775. DOI.
- Newman S, Nakitandwe J, Kesserwan CA, et al. Genomes for Kids: The Scope of Pathogenic Mutations in Pediatric Cancer Revealed by Comprehensive DNA and RNA Sequencing. Cancer Discovery. 2021;11:3008-3027. DOI.
- Oka M, Xu L, Suzuki T, et al. Aberrant splicing isoforms detected by full-length transcriptome sequencing as transcripts of potential neoantigens in non-small cell lung cancer. Genome Biology. 2021;22:9. DOI.
- Ouspenskaia T, Law T, Clauser KR, et al. Unannotated proteins expand the MHC-I-restricted immunopeptidome in cancer. Nature Biotechnology. 2022;40:209-217. DOI.
- Ely ZA, Kulstad ZJ, Gunaydin G, et al. Pancreatic cancer-restricted cryptic antigens are targets for T cell recognition. Science. 2025;388:eadk3487. DOI.
More sequencing resources and application guides are available in the Biomedical NGS Learning Center.
Notes on sources and scope: This article summarizes independent published studies to illustrate how different genomic, transcriptomic, and translation-level evidence can support neoantigen discovery. Published study findings are not service performance data. Sequencing strategy should be selected according to tumor biology, specimen availability, and the downstream research question.