What Samples and Files Are Needed to Start a Personalized Neoantigen Project?

Personalized neoantigen project input framework showing tumor, matched normal, RNA, HLA evidence, metadata, and sequencing filesFigure 1. A project-ready input package connects biological specimens, matched sequencing data, HLA evidence, and traceable metadata.

A personalized neoantigen project is not started by uploading a tumor variant list alone. Reliable candidate generation depends on a connected evidence set: tumor DNA, matched-normal DNA, tumor RNA, an appropriate HLA genotype, sample and library metadata, reference versions, and files that preserve quality and provenance. The exact package can vary with the biological question and available material, but missing normal, expression, purity, or HLA information changes what the analysis can claim. This guide provides a practical intake framework for research teams preparing new specimens, transferring existing sequencing files, or combining both.

CD Genomics supports sequencing and bioinformatics for research use. These services are not intended for clinical diagnosis or individual treatment selection.

Key Takeaways

  • Start with a matched sample map. Every tumor specimen should be linked unambiguously to its normal comparator, RNA aliquot, time point, and participant-level identifier.
  • Preserve raw evidence. FASTQ files plus complete metadata generally support more consistent reanalysis than filtered variant tables alone.
  • Treat RNA as functional prioritization evidence. Tumor RNA can confirm gene and allele expression, reveal aberrant transcripts, and identify quality problems that DNA cannot resolve.
  • Make HLA provenance explicit. Record whether HLA alleles came from dedicated typing, DNA inference, RNA inference, or a prior report, together with the caller and database version.
  • Agree on references and deliverables before transfer. Genome build, transcript annotation, naming conventions, filters, and output fields should be locked at project kickoff.

Define the Research Question First

The input package should follow the intended endpoint. A discovery project may seek a broad, auditable list of expressed somatic candidates. A longitudinal project may compare candidates across time points. A functional study may require peptide sequences and supporting evidence for downstream assays. These goals change minimum sample numbers, depth expectations, replicate design, and reporting rules.

Before preparing material, document whether the analysis should include single-nucleotide variants, small insertions and deletions, gene fusions, splice-derived events, frameshifts, or other candidate sources. RNA splicing can add candidate classes beyond conventional coding DNA variants, but it requires suitable RNA data and dedicated filters [5]. A broad scope without matching assays usually produces an uneven evidence set. A recent renal cell carcinoma neoantigen-vaccine study illustrates how candidate nomination can be linked to tumor sequencing, HLA context, manufacturing constraints, and downstream immune measurements in a defined research program [3].

The core data flow commonly combines Whole Exome Sequencing for tumor-normal coding variants, Transcriptome Sequencing for expression and transcript evidence, and HLA Typing by Sequencing or validated HLA calls for peptide-presentation modeling. The project brief should state which source is authoritative when different assays disagree.

Research objective Minimum evidence Valuable addition Main limitation if absent
Expressed coding candidates Tumor DNA, matched normal, tumor RNA, HLA genotype Replicate RNA or orthogonal variant review DNA-only candidates may not be expressed
Indel or frameshift candidates High-quality paired DNA data and transcript annotation RNA support spanning the altered transcript Alignment and annotation artifacts can dominate
Fusion-derived candidates Tumor RNA with adequate library design Tumor DNA structural evidence Caller agreement and reading-frame uncertainty
Splice-derived candidates Stranded or well-documented RNA-seq and normal-reference strategy Long-read transcript evidence Aberrant junctions may reflect technical noise
Longitudinal comparison Consistently processed tumor time points and one matched normal Repeated RNA and immune profiling Batch effects can mimic biological change
Functional follow-up Sequence-level candidate evidence and HLA context TCR or immunopeptidomic evidence Prediction rank alone does not establish recognition

Assemble the Biological Sample Set

The preferred foundation is paired tumor and germline material from the same research participant. The normal comparator helps remove inherited variation and sequencing artifacts. Blood-derived DNA is common, but its suitability depends on the study context; an alternate uninvolved tissue may be preferable when blood can contain disease-related or age-associated clones. Whatever is used, record the tissue source and collection time.

Tumor DNA requirements

Tumor specimens may be fresh frozen, stabilized, or formalin-fixed paraffin-embedded. Record estimated tumor content, necrosis, input amount, extraction method, storage history, block age when applicable, and any prior molecular processing. Low tumor purity reduces somatic variant allele fractions. Degraded DNA changes insert sizes and can increase artifacts, especially in fixed material.

Macrodissection or another enrichment strategy may improve tumor fraction, but pre- and post-enrichment identifiers must remain connected. If several regions are available, decide whether they will be analyzed separately to study heterogeneity or pooled to obtain one composite profile. Combining regions before sequencing removes spatial information.

Matched-normal DNA requirements

Matched-normal identity is as important as tumor quality. Use a unique participant identifier that is distinct from sample and library identifiers. A swap between tumor and normal can create thousands of apparent somatic variants, so the analysis should include genetic concordance and contamination checks.

When no matched normal is available, a tumor-only workflow can use population databases and panels of normal samples. However, residual germline variants and technical artifacts generally increase. The report must label the design as tumor-only and avoid presenting its candidate set as equivalent to paired analysis.

Tumor RNA requirements

RNA should ideally originate from the same tumor region or a closely matched aliquot. Record extraction method, quantity, integrity metric, library type, strandedness, capture or depletion method, read layout, and batch. Fixed samples may require DV200 or another degradation-aware metric in addition to RIN.

RNA supports expression filtering, allele-specific read review, transcript selection, fusion discovery, and splice-event assessment. A Kenyan breast cancer study integrating exome and RNA sequencing illustrates how these layers produce sample-specific neoantigen profiles [2]. RNA does not replace tumor DNA for confident somatic calling: transcriptional editing, mapping ambiguity, allele-specific expression, and variable coverage can complicate RNA-only variant detection.

Sample decision tree linking tumor material, matched normal, RNA quality, and optional longitudinal or immune samplesFigure 2. Sample selection should preserve matching, tumor context, and the evidence needed for each candidate class.

Add HLA and Optional Immune Evidence

Peptide-binding predictions depend on the submitted HLA genotype. At minimum, provide class I alleles for HLA-A, HLA-B, and HLA-C at unambiguous allele resolution appropriate to the chosen prediction tool. Add class II loci when CD4-associated candidates are in scope. HLA-unbiased genetic screens have recovered patient-specific CD4 as well as CD8 neoantigens, so class II should not be treated as a decorative field when it is part of the research question [6].

Dedicated typing data can reduce ambiguity, while inference from exome or RNA sequencing may be adequate in some designs. The HLA study design guide explains when locus coverage, phasing, and ambiguity matter. The HLA platform choice guide can help align resolution with project goals.

Submit allele calls exactly as reported and include the typing method, software, reference database release, quality flags, and alternative calls. Do not silently truncate four-field alleles or convert them to two-field notation. If tumor HLA loss of heterozygosity is evaluated, retain tumor-normal copy-number inputs and purity estimates; computational sensitivity varies with clonality and coverage [4].

Optional immune evidence may include paired or bulk TCR repertoires, peptide-screen results, immunopeptidomics, or existing functional assay measurements. These inputs can strengthen prioritization but should remain labeled by evidence type. The BCR and TCR Sequencing service can support repertoire-focused research when a study includes clonotype abundance or tracking.

Prepare Sequencing and Analysis Files

Raw paired-end FASTQ files are the most portable starting point for reanalysis. If raw files cannot be transferred, coordinate acceptance of aligned BAM or CRAM files before the project begins. Alignment files need headers, read groups, sample names, reference contigs, and indexes; CRAM also requires the exact reference sequence used for compression.

For each tumor-normal DNA pair, provide:

  • R1 and R2 FASTQ files or complete BAM/CRAM plus index;
  • sample, library, lane, platform, read-length, and capture-kit identifiers;
  • genome build and decoy or alternate-contig details;
  • duplicate-marking, base-quality processing, and alignment software versions if preprocessed;
  • any existing somatic VCF, copy-number, purity, contamination, or concordance results;
  • the capture target BED file and vendor design version for exome data.

For each RNA sample, provide:

  • paired FASTQ files or a coordinate-sorted alignment with index;
  • library protocol, strandedness, poly(A) selection or rRNA depletion, and read layout;
  • genome build, transcript annotation release, and alignment or quantification version;
  • existing gene-expression matrices with units clearly stated;
  • fusion or splice-junction results together with caller and filtering parameters;
  • RNA quality metrics and the relationship to the paired tumor DNA aliquot.

Previously generated VCF files can accelerate review, but they are not self-explanatory. Include the header and annotation fields, caller version, tumor and normal sample columns, genome build, filtering status, normalization method, and whether multiallelic records were split. A spreadsheet containing chromosome, position, and amino-acid change alone is usually insufficient to reproduce peptide generation.

Harmonize references before reuse

Reusing public or previously processed data requires an explicit harmonization step. Confirm whether coordinates use GRCh37, GRCh38, or another assembly; whether chromosome prefixes and mitochondrial contigs match; and whether transcript identifiers belong to the expected Ensembl, RefSeq, or custom annotation release. Coordinate liftover can move a variant, but it does not automatically reproduce the original transcript consequence or repair alleles that no longer match the destination reference.

If DNA and RNA were processed against different builds, retain the original alignments and derived files, then agree whether one assay will be reprocessed. Likewise, do not merge expression values produced as raw counts, TPM, FPKM, and transformed abundance without a documented conversion or analysis plan. Capture kits, read length, and sequencing batches should remain available as covariates when samples come from several sources.

File acceptance should include spot checks that variant reference alleles match the declared genome, BAM or CRAM contigs match the indexes, transcript identifiers resolve in the chosen annotation, and HLA allele names are valid in the relevant database release. These inexpensive checks detect incompatible inputs before peptide generation propagates the errors.

File-manifest workflow showing FASTQ, BAM or CRAM, VCF, expression, HLA, metadata, references, and checksumsFigure 3. A transfer-ready file package retains raw data, derived evidence, software provenance, reference versions, and checksums.

Use a Manifest That Prevents Silent Mismatches

Create one row per biological specimen and separate columns for participant, specimen, extraction, library, and file identifiers. Reusing one short name for all five levels makes failures hard to trace. The manifest should include pairing, tissue, collection time, preservation, extraction, input amount, QC, tumor content, assay, library batch, sequencing run, file paths, and checksums.

Manifest field Example purpose Validation rule
participant_id Links tumor and normal without exposing direct identifiers One stable pseudonymous value per participant
specimen_id Distinguishes tissue collections and time points Unique across the project
analyte_id Separates DNA and RNA extractions Maps to exactly one specimen
library_id Tracks preparation batch and assay Maps to exactly one analyte
role Labels tumor, normal, or other comparator Controlled vocabulary
pair_id Connects analysis-ready tumor-normal pairs No orphan tumor unless approved
genome_build Prevents coordinate mismatch Exact release, not "human genome"
hla_source Records typing or inference provenance Includes method and version
file_checksum Detects transfer corruption Recalculated after receipt
consent_scope_code Encodes allowed research handling Reviewed before processing

Use pseudonymous research identifiers in filenames and manifests. Transfer only the metadata required for the agreed analysis, through an approved secure channel. The scientific intake sheet and the data-governance record can be linked by a controlled project identifier without placing direct personal identifiers in analysis filenames.

Lock the Computational Contract

Two teams can analyze the same files and return different candidate lists because of differences in genome build, transcript selection, somatic filters, HLA calls, peptide lengths, binding predictors, expression thresholds, and ranking logic. The project specification should capture these choices before the final run.

Define how candidates will be generated from missense variants, frameshifts, fusions, and splice events; how phased variants will be handled; which transcripts are eligible; what counts as expressed; and how ambiguous HLA calls enter prediction. State whether wild-type peptide comparison, variant allele expression, clonality, HLA loss, manufacturability, or functional evidence contributes to ranking.

HLA-unbiased genetic screens have demonstrated that patient-specific CD4 and CD8 targets can be identified without relying solely on prediction [6]. Personalized TCR-replacement research has also shown the scale of integrated mutation, HLA, screening, and receptor evidence required to move from candidates to experimentally supported targets [1]. These studies reinforce a practical distinction: computational ranking creates hypotheses, while recognition evidence tests them.

The immuno-oncology research solutions page provides context for combining genomic, transcriptomic, and immune-repertoire assays. A kickoff should convert that broad capability into a fixed sample-to-deliverable matrix for the specific project.

Record decisions and exceptions

Maintain a decision log for every deviation from the locked workflow. Examples include accepting a tumor-only sample, changing an expression threshold, selecting one of several HLA calls, excluding a contaminated library, or substituting a transcript. Each entry should identify the affected sample, evidence reviewed, decision owner, date, and downstream fields that need to carry a limitation flag. This prevents later reruns from silently applying different rules to similar samples.

Run a Preflight Check Before Transfer

A short preflight review prevents expensive rework. Confirm sample identity, pairing, input QC, file readability, checksum agreement, build compatibility, HLA completeness, and metadata coverage before sequencing or analysis begins.

Recommended acceptance checks include:

  • tumor-normal genetic concordance and sex-chromosome consistency where appropriate;
  • contamination estimates and expected tumor purity range;
  • DNA library insert size, duplication, on-target performance, and coverage distribution;
  • RNA mapping, strandedness, ribosomal content, exonic distribution, and expression complexity;
  • consistency among BAM headers, VCF sample columns, filenames, and the manifest;
  • matching reference builds for variants, alignments, targets, and annotations;
  • valid checksums and a documented replacement path for corrupted files;
  • HLA nomenclature validation and explicit resolution of ambiguous calls.

The ambiguous HLA calls guide describes how additional data or qualified reporting can handle unresolved alleles. Do not choose an allele merely because it produces a more favorable prediction.

Neoantigen project preflight dashboard showing identity, sequencing quality, genome-build compatibility, HLA completeness, and metadata readinessFigure 4. Preflight review separates ready inputs from correctable gaps and limitations that must remain in the report.

Define the Deliverable Package

A useful final package connects each candidate to its source evidence. Request a candidate table containing variant or event identifiers, transcript and protein consequences, mutant and wild-type peptide sequences, peptide length, HLA allele, prediction outputs, DNA support, RNA support, expression context, clonality-related fields when available, filters, and rank rationale.

Also request sample and software manifests, QC summaries, intermediate variant and HLA files, and machine-readable tables. A ranked PDF alone is difficult to audit or reanalyze. The ESCAPE-seq neoantigen research overview offers an example of connecting candidate discovery with experimental prioritization, while the TMB and HLA diversity guide explains why mutation burden and HLA context should be interpreted as separate dimensions.

FAQ

  • Is matched-normal DNA mandatory?
  • Can existing BAM and VCF files replace FASTQ files?
  • Is tumor RNA required?
  • What HLA information should be supplied?
  • What is the most common avoidable intake problem?

References

  1. Foy SP, Jacoby K, Bota DA, Hunter T, Pan Z, Stawiski E, Ma Y, Lu W, Peng S, Wang CL, Yuen B, Dalmas O, Heeringa K, Sennino B, Conroy A, Bethune MT, Mende I, White W, Kukreja M, Gunturu S, Humphrey E, Hussaini A, An D, Litterman AJ, Quach BB, Ng AHC, Lu Y, Smith C, Campbell KM, Anaya D, Skrdlant L, Huang EYH, Mendoza V, Mathur J, Dengler L, Purandare B, Moot R, Yi MC, Funke R, Sibley A, Stallings-Schmitt T, Oh DY, Chmielowski B, Abedi M, Yuan Y, Sosman JA, Lee SM, Schoenfeld AJ, Baltimore D, Heath JR, Franzusoff A, Ribas A, Rao AV, Mandl SJ. Non-viral precision T cell receptor replacement for personalized cell therapy. Nature. 2023;615(7953):687-696. doi:10.1038/s41586-022-05531-1
  2. Wagutu G, Gitau J, Mwangi K, Murithi M, Melly E, Harris AR, Sayed S, Ambs S, Makokha F. Whole exome-seq and RNA-seq data reveal unique neoantigen profiles in Kenyan breast cancer patients. Frontiers in Oncology. 2024;14:1444327. doi:10.3389/fonc.2024.1444327
  3. Braun DA, Moranzoni G, Chea V, et al. A neoantigen vaccine generates antitumour immunity in renal cell carcinoma. Nature. 2025;639(8054):474-482. doi:10.1038/s41586-024-08507-5
  4. Pyke RM, Mellacheruvu D, Dea S, Abbott CW, McDaniel L, Bhave DP, Zhang SV, Levy E, Bartha G, West J, Snyder MP, Chen RO, Boyle SM. A machine learning algorithm with subclonal sensitivity reveals widespread pan-cancer human leukocyte antigen loss of heterozygosity. Nature Communications. 2022;13(1):1925. doi:10.1038/s41467-022-29203-w
  5. Wickland DP, McNinch C, Jessen E, Necela B, Shreeder B, Lin Y, Knutson KL, Asmann YW. Comprehensive profiling of cancer neoantigens from aberrant RNA splicing. Journal for ImmunoTherapy of Cancer. 2024;12(5):e008988. doi:10.1136/jitc-2024-008988
  6. Cattaneo CM, Battaglia T, Urbanus J, Moravec Z, Voogd R, de Groot R, Hartemink KJ, Haanen JBAG, Voest EE, Schumacher TN, Scheper W. Identification of patient-specific CD4+ and CD8+ T cell neoantigens through HLA-unbiased genetic screens. Nature Biotechnology. 2023;41(6):783-787. doi:10.1038/s41587-022-01547-0

For research use only. Not for use in diagnostic procedures or individual treatment decisions.

For research purposes only, not intended for clinical diagnosis, treatment, or individual health assessments.


Related Services
Inquiry
For research purposes only, not intended for clinical diagnosis, treatment, or individual health assessments.