TCR and BCR Immune Repertoire Sequencing: Profiling T-Cell and B-Cell Receptor Diversity in Health and Disease
Meta Intent: A practical research-design guide for deciding what a TCR or BCR repertoire result can credibly claim, then selecting the sample architecture, assay, clonotype definition, and analysis plan needed to support that claim.
Immune repertoire sequencing is often described as a way to measure diversity, detect expanded clones, and track adaptive immunity. Those descriptions are useful, but they skip the decision that determines whether the final result will be persuasive: what exactly must be observed for the biological conclusion to hold? A ranked clonotype table can show that sequences were recovered. It cannot, by itself, establish antigen specificity, a stable cell lineage, a meaningful longitudinal change, or a functional receptor. This resource starts with the evidence claim and works backward through assay choice, sampling, and analysis. It is deliberately different from the linked HLA articles in this content matrix: the focus here is the integrity of clonal evidence, not allele-resolution reporting or donor-unit matching.
A Repertoire Is a Sampled Molecular Census, Not a Direct Antigen Map
Every repertoire dataset is a census of receptor molecules or rearrangements recovered from a finite aliquot. The unit being counted may be a genomic rearrangement, an expressed transcript, a UMI-collapsed molecule, a cell barcode, or a paired receptor carried by one cell. These are not interchangeable. A clone that is abundant in RNA may be transcriptionally active, represented by more cells, or both. A clone absent from a later sample may be biologically absent, below the sampling limit, lost before library construction, or filtered by an analysis rule.
Start each project with one sentence in the protocol: "We will use this assay to estimate [defined receptor feature] in [defined cellular compartment] at [defined comparison point], and will interpret a difference only after [defined technical checks]." That sentence prevents the common error of upgrading an observation into a stronger claim. For example, "expanded TRB CDR3 sequences in sorted memory T cells" is an observation; "a persistent antigen-specific T-cell response" is a hypothesis that needs additional evidence. The same distinction matters for B cells: shared IGH sequences may establish related rearrangements, while a clonal family with a coherent somatic-hypermutation pattern needs an explicitly stated lineage model.
Figure 1: The Repertoire Evidence Chain. A result becomes more biologically specific only when its definition, sampling frame, and orthogonal evidence support the next step in the evidence chain.
Choose the Quantification Unit Before Choosing RNA, DNA, or Cell Barcodes
The first assay decision is not "TCR or BCR?" It is "what must the abundance represent?" Genomic DNA starts close to a cell-count model because each lymphocyte contributes rearranged receptor loci. It is useful when a study needs a durable record of rearrangement abundance or must work with DNA that is more stable than RNA. RNA captures expressed receptor transcripts, supports full variable-region information more naturally, and is especially valuable when BCR somatic hypermutation, isotype context, or transcript-level receptor features are central. But RNA abundance should not be presented as a direct cell count without appropriate molecular correction and a clear statement of the measurement unit.
For RNA-based designs, molecular identifiers and replicate-aware analysis help separate molecules from PCR duplicates. They do not remove every source of variance: cell capture, RNA extraction, reverse transcription, and initial aliquoting still define which molecules were available to observe. A project using TCR and BCR immune repertoire sequencing should therefore record source material, cell-enrichment history, nucleic-acid type, library chemistry, and the counting unit in a machine-readable project sheet. When the research question also requires broader expression context from the same material, an total RNA sequencing design can provide a complementary transcript-level layer; it should not be treated as a substitute for a receptor-enriched repertoire assay.
A practical rule is to avoid a silent switch in denominator. If one visit is reported as "fraction of UMI-collapsed receptor molecules" and another as "fraction of productive clonotypes," the resulting percentages cannot be compared as a longitudinal trajectory. Keep a stable primary denominator, retain secondary denominators for sensitivity analysis, and document any sample where the denominator changes because of low input or a modified workflow.
Bulk and Single-Cell Repertoire Sequencing Answer Different Questions
Bulk repertoire sequencing is well suited to broad, repeatable surveillance of a defined cell population across many samples or time points. It can recover a long tail of receptor sequences from a pooled population and supports questions such as: Which clonotypes are repeatedly detected? Has the frequency distribution become more concentrated? Are V/J usage or CDR3 properties shifted between prespecified groups? In this mode, the main design challenge is representativeness: the cells entering the tube, the nucleic acid entering the library, and the molecular depth entering the analysis all constrain the claim.
Single-cell receptor profiling should be selected when chain pairing or cellular context changes the biological interpretation. A paired alpha-beta TCR or heavy-light BCR sequence is not merely more detailed than one chain; it is a different evidence object. It can be linked to a cell state, a surface-marker panel, a transcriptional program, or a clone-specific phenotype. The appropriate bridge is a single-cell TCR and BCR sequencing service when the project genuinely needs paired receptors or cell-level context. Do not use single-cell profiling simply because it is more granular if the central question is a cohort-scale change in one unpaired chain across serial samples. In that case, missing low-frequency clones and lower per-sample repertoire coverage may weaken the primary endpoint.
Many strong studies use a staged design rather than forcing one assay to do both jobs. A bulk screen identifies time points, tissue compartments, or samples with a prespecified repertoire signal. A focused single-cell cohort then tests whether the candidate signal is associated with a paired receptor and a defined cellular state. This sequence makes the escalation rule explicit and keeps the paired-cell analysis tied to a question that bulk data cannot answer. When transcriptome context is the intended second layer, RNA-Seq can be planned as a complementary expression dataset, with sample identifiers and collection windows harmonized before either library is generated.
Figure 2: Assay Selection by Evidence Requirement. Assay selection should follow the evidence object needed for the final conclusion: population-level frequency, paired receptor identity, or receptor identity linked to cell state.
Define What Counts as the Same Clone Before Looking at the Data
"Clonotype" is not a universal object. In a bulk TCR dataset, a study may define a clonotype by productive CDR3 nucleotide sequence plus V and J assignment for one chain. Another may use amino-acid CDR3 plus gene calls to group related sequences. In single-cell data, a clonotype may be a paired alpha-beta receptor configuration; in BCR analysis, a clonal family may incorporate V and J assignment, junction similarity, and a model of somatic mutation. Each definition has consequences for clone counts, sharing, overlap, and diversity values.
Write the definition before reviewing group differences. At minimum, define whether sequences must be productive, whether the identity key is nucleotide or amino acid sequence, which gene-call ambiguity rules are accepted, how multiple chains in one cell are handled, and how unpaired contigs are retained or excluded. Store this definition beside the data as fields such as clone_definition, productive_rule, reference_release, and contig_filter_version. This is more useful than a generic statement that a pipeline was "standard." It lets a future analyst reproduce why two sequences were grouped together.
For BCR projects, the definition should separate three levels: a unique rearrangement, a related clonal family, and a candidate lineage path. Somatic hypermutation makes these levels especially easy to blur. An observed family of related IGH sequences is not automatically a directional maturation trajectory, and a sequence resemblance is not proof of shared antigen binding. If full variable-region context is required for the research question, full-length TCR and BCR repertoire profiling gives a more appropriate data object than a CDR3-only claim. If a selected antibody sequence later needs independent molecular characterization, that is a separate task from population-level repertoire profiling and can be scoped through antibody sequencing.
Build Sampling Around the Biological Signal, Not Around a Convenient Tube
The largest source of uncertainty in repertoire work frequently occurs before sequencing. Peripheral blood, tissue, a sorted subset, and a cell-enriched suspension are different biological sampling frames. A receptor sequence can be common in a lesion yet uncommon in circulation; a dominant sequence in a bulk tissue specimen may originate from a small cellular compartment. The protocol should state which compartment represents the question and whether comparisons cross compartments. If they do, the comparison must be framed as trafficking, sharing, or enrichment—not as a direct equivalence of frequencies.
For longitudinal studies, standardize collection windows relative to the research event, preservation method, processing delay, enrichment strategy, and the cell population entering each library. A "baseline" sample collected after an unrecorded perturbation is not baseline. Likewise, a shift from whole PBMCs to sorted CD8-positive cells changes the denominator and should be analyzed as a new measurement series unless a planned bridge sample establishes comparability. Where the study needs a small, prespecified set of genomic targets as an orthogonal check, targeted region sequencing can be designed separately; it should validate a defined target or sample identity rather than be used to infer an entire receptor repertoire.
Technical replicates are most informative when they answer a known uncertainty. Replicate libraries from the same nucleic-acid extract test library and sequencing variability; split cell aliquots test an earlier sampling stage; independent collections test biological variability. Label these roles clearly. Combining them into one "replicate" column hides the exact variation that a later sensitivity analysis needs to evaluate.
Figure 3: Sampling Architecture for Longitudinal Repertoire Studies. Separating biological collection, cell selection, nucleic-acid aliquoting, and library replication reveals where a reported change could have arisen.
Use Coverage Diagnostics Before Interpreting Diversity or Sharing
Diversity metrics are attractive because they turn a complex repertoire into a single number. Their weakness is that the number is inseparable from the observed sample. Shannon entropy, Simpson diversity, clonality, Gini-based concentration measures, top-clone share, richness, and overlap indices emphasize different parts of the clone-frequency distribution. A result can therefore change because a few dominant clones expand, because rare clones are missed, because the cell population shifts, or because the detection threshold changed.
Make coverage diagnostics part of the primary report. Rarefaction curves, read- or UMI-subsampling, productive-sequence fractions, duplicate burden, and the distribution of clone support show whether the low-frequency tail is still being discovered. A rank-abundance plot should be accompanied by its denominator and a statement of whether values were normalized before cross-sample comparison. For a project comparing groups with unequal usable depth, predefine a common resampling strategy or report depth-stratified sensitivity results. Downsampling does not restore clones that never entered the sample; it does prevent a deeper library from appearing more diverse merely because it had more opportunity to detect rare sequences.
Choose one primary metric that corresponds to the hypothesis, then a small set of diagnostic metrics that challenge it. If the hypothesis concerns dominant clonal expansion, a concentration measure and a top-clone trajectory may be primary; richness can be supportive but should not carry the conclusion. If the hypothesis concerns repertoire breadth, rarefaction and low-frequency robustness become central. CD Genomics can align these choices with the intended output of TCR and BCR immune repertoire sequencing so that the exported clonotype table, metric definitions, and requested visualizations refer to the same analysis contract.
Figure 4: Coverage Diagnostics Before Diversity Metrics. Rarefaction, molecular support, and rank-abundance diagnostics reveal whether a diversity or sharing comparison is driven by coverage rather than biology.
Interpret Longitudinal Clonotype Tracking as a Recapture Problem
Serial repertoire studies are commonly summarized as "appeared," "persisted," or "disappeared." Those words imply a certainty that most assays do not provide. A better framing is recapture probability: given the clone's observed support at the first time point, the sample size and diversity at the next time point, and the technical threshold, how likely was it to be observed again if it remained present? Large, well-supported clonotypes are more likely to be recaptured than singletons. This does not make small clones irrelevant; it means their non-detection requires more cautious language.
Predefine the evidence categories. For example, "consistently observed" can require detection under the same clone definition in more than one planned collection, with a stated minimum molecular support. "Not detected" should remain distinct from "lost." "Expansion" should require a change that exceeds replicate-informed technical variation and is visible under the same normalization rule. A longitudinal report should include a clone-level trace for the prespecified candidates, a population-level distribution summary, and a record of samples that failed coverage criteria rather than silently excluding them.
For research programs where receptor dynamics are interpreted beside inherited immune-genetic context, connect the two layers explicitly rather than conflating them. HLA typing describes germline allele information; repertoire sequencing describes somatically rearranged populations. The matrix hub, HLA Typing and Immune Repertoire Sequencing Services, explains how these evidence types can be planned together. The companion guide, High-Resolution HLA Typing by NGS, should be used when the question concerns the reportable resolution of the HLA call rather than clonal dynamics.
Figure 5: From Detection to Longitudinal Recapture. Longitudinal interpretation should account for the fact that recapture probability depends on initial clone support and the effective sampling depth at each time point.
Keep BCR Lineage Inference Separate From Receptor Function Claims
BCR datasets provide a second layer of complexity because related sequences can accumulate somatic mutations. A rigorous BCR analysis first assigns V, D, and J segments against a stated reference release, then applies a documented rule for grouping sequences into candidate clonal families. Only after that should lineage reconstruction, mutation-spectrum analysis, isotype patterns, or shared-family comparisons be considered. Each step should retain the uncertainty introduced by sequencing errors, amplification artifacts, incomplete sequence coverage, and germline-reference ambiguity.
A tree-like diagram is a useful exploration tool, but it is not a functional assay. It can support a statement such as "these sequences are consistent with a related B-cell clonal family under the stated grouping rule." It cannot establish antigen binding, affinity, neutralization, or biological activity without orthogonal experiments. This limit is productive: it tells the team what to validate next. When a project needs receptor sequence, cell state, chromatin accessibility, and expression interpreted together, a planned multi-omics service can organize the relevant layers. An ATAC-Seq dataset may be valuable when the question is whether clonally related cells occupy distinct regulatory states, not when the only needed endpoint is a bulk clone-frequency shift.
Figure 6: BCR Clonal Families and the Limits of Lineage Inference. BCR family grouping and lineage reconstruction can organize related sequences, but receptor function requires evidence beyond sequence similarity.
Design an Analysis Contract That Can Survive Reanalysis
Reanalysis is expected in repertoire work. Germline databases update, annotation algorithms improve, productive-contig rules are refined, and collaborators may ask a different question of the same material. A defensible project retains the raw reads, sample manifest, processing parameters, reference versions, and intermediate annotation outputs needed to regenerate the final table. The deliverable should not be only a PDF with plots.
Before sequencing, define a compact analysis contract containing: the primary biological question; population and compartment; nucleic-acid and assay choice; required chains; clone identity rule; primary denominator; minimum support rule; coverage diagnostics; primary and sensitivity metrics; longitudinal comparison logic; and exclusions. Add identifiers such as participant_id, collection_window, cell_population, library_id, umi_policy, and analysis_version to the manifest. These fields make a future merge possible without reconstructing study history from email threads.
This approach also clarifies when an additional assay is justified. Escalate from bulk to paired single-cell sequencing when pair identity or phenotype is required for the claim. Escalate to full-length receptor coverage when variable-region context changes BCR or TCR interpretation. Escalate to multi-omic context when a clone-frequency result needs a state-level explanation. Do not escalate merely to create a larger data package. Each new layer should close a named evidence gap.
Use Three Predeclared Go/No-Go Checks
- Coverage: Do rarefaction and molecular-support diagnostics meet the condition declared before analysis?
- Definition stability: Does the conclusion persist under the stated clone definition and normalization choice?
- Evidence boundary: Is the conclusion no stronger than the molecular evidence measured by this assay?
Figure 7: The Repertoire Analysis Contract. A versioned analysis contract connects the research question to the clone definition, quality gates, and reportable conclusion.
What a Decision-Ready Repertoire Deliverable Should Contain
A useful final package has two levels. The first is a readable interpretation layer: sample and library QC, clone-definition language, coverage diagnostics, prespecified metrics, and visualizations matched to the research question. The second is a reusable data layer: raw or demultiplexed sequence files as appropriate, annotated rearrangement records, clonotype and clonal-family tables, a sample manifest, versioned analysis settings, and a data dictionary. The data layer is what allows an investigator to test a new comparison without changing the original meaning of each field.
For a bulk project, ask whether the report distinguishes reads, molecules, and clonotypes; whether productive and nonproductive sequences are separately handled; whether rarefaction or comparable coverage diagnostics are included; and whether all normalization steps are explicit. For a single-cell project, ask whether paired and unpaired chains are separately reported, how cells with multiple contigs were handled, and how receptor calls were linked to cell annotations. For BCR work, ask whether germline assignment, family clustering, mutation analysis, and lineage inference are clearly separated. These questions make a service conversation more efficient because they define the data objects required before the project begins.
How CD Genomics Supports Question-Led Immune Repertoire Studies
CD Genomics supports research teams that need to turn an immune-repertoire question into an assay and analysis plan with explicit evidence boundaries. A project discussion can begin with the required receptor chains, source material, comparison structure, and expected analysis objects rather than with a generic request for "deep sequencing." The goal is a report that can be read alongside its clonotype tables, metadata, and quality controls—not a visually polished result that cannot be reinterpreted.
For studies centered on serial population-level dynamics, begin with the primary TCR and BCR immune repertoire sequencing service. For projects that need paired receptors and cellular context, use the single-cell TCR and BCR sequencing service as the appropriate escalation. Each escalation should close a stated evidence gap, not merely enlarge the data package.
Frequently Asked Questions
Can bulk TCR sequencing identify paired alpha and beta chains?
No. Bulk data can characterize one or more chains in a pooled population, but it does not preserve which alpha and beta chains came from the same cell. Use a paired single-cell design when that relationship is required.
Is a clone that is not detected at the next time point gone?
Not necessarily. Non-detection can reflect biological loss or incomplete recapture. Interpret it with the clone's earlier support, effective depth, coverage diagnostics, and the prespecified detection rule.
Should RNA or DNA be used for repertoire sequencing?
Choose RNA when expressed receptor features and full variable-region context are central. Choose DNA when the research question needs a rearrangement-oriented abundance measure or RNA quality is limiting. State the resulting measurement unit in the analysis contract.
Does a shared CDR3 sequence prove shared antigen specificity?
No. Sequence sharing can prioritize a hypothesis, but antigen specificity needs independent evidence. The required validation depends on the receptor, model, and proposed biological claim.
Which diversity metric is best?
There is no universal best metric. Select one primary metric that matches the hypothesis, then use coverage diagnostics and complementary metrics to test whether the conclusion is stable.
When should BCR projects include lineage analysis?
Include it when related-sequence structure or somatic-mutation patterns are part of the question and sufficient variable-region information is available. Keep lineage inference distinct from functional antibody claims.
Research Use Notice: This resource is for research use only.
References:
- Ma K-Y, et al. Immune Repertoire Sequencing Using Molecular Identifiers Enables Accurate Clonality Discovery and Clone Size Quantification. Frontiers in Immunology. 2018. DOI: 10.3389/fimmu.2018.00033
- De Simone M, Rossetti G, Pagani M. Single Cell T Cell Receptor Sequencing: Techniques and Future Challenges. Frontiers in Immunology. 2018. DOI: 10.3389/fimmu.2018.01638
- López-Santibáñez-Jácome L, Avendaño-Vázquez SE, Flores-Jasso CF. The Pipeline Repertoire for Ig-Seq Analysis. Frontiers in Immunology. 2019. DOI: 10.3389/fimmu.2019.00899
- de Greef PC, et al. On the Feasibility of Using TCR Sequencing to Follow a Vaccination Response: Lessons Learned. Frontiers in Immunology. 2023. DOI: 10.3389/fimmu.2023.1210168
- Pavlova AV, Zvyagin IV, Shugay M. Detecting T-Cell Clonal Expansions and Quantifying Clone Survival Using Deep Profiling of Immune Repertoires. Frontiers in Immunology. 2024. DOI: 10.3389/fimmu.2024.1321603
- Aoki T, Shichino S, Matsushima K, Ueha S. Revealing Clonal Responses of Tumor-Reactive T-Cells Through T Cell Receptor Repertoire Analysis. Frontiers in Immunology. 2022. DOI: 10.3389/fimmu.2022.807696
Related Services