Single-Cell V(D)J and CITE-Seq: Integrating Immune Repertoire Profiling and Multi-Omics at Single-Cell Resolution

A single droplet in a 10x Chromium chip can now capture four layers of information from one cell: the full transcriptome, paired TCR or BCR sequences, surface protein abundance for dozens of markers, and antigen-binding specificity — all indexed to the same cell barcode. This convergence of modalities, enabled by the integration of V(D)J sequencing and CITE-Seq (Cellular Indexing of Transcriptomes and Epitopes by Sequencing) into the 10x Genomics Single Cell Immune Profiling platform, has transformed how immunologists interrogate adaptive immunity. A T cell is no longer defined by its transcriptome alone — it is defined by the clonotype it expresses, the cytokine receptors on its surface, the transcription factors in its nucleus, and, when BEAM (Barcode Enabled Antigen Mapping) is layered on, the peptide-MHC complex it recognizes. This guide covers the V(D)J and CITE-Seq workflows, from sample preparation and panel design through data integration and biological interpretation.

V(D)J and CITE-Seq Multimodal WorkflowFigure 1: V(D)J + CITE-Seq Multimodal Workflow — From Single-Cell Suspension to Four-Layer Data Output

V(D)J Sequencing — Decoding the Adaptive Immune Repertoire

The adaptive immune system stores its memory in DNA. During lymphocyte development, T and B cells undergo V(D)J recombination — a site-specific somatic recombination process that assembles variable (V), diversity (D), and joining (J) gene segments into a functional antigen receptor gene. The combinatorial diversity generated by this process, amplified by junctional diversification (nucleotide deletion and N-nucleotide addition at segment junctions), produces a theoretical repertoire exceeding 10^15 unique receptors, dwarfing the roughly 10^12 lymphocytes in the human body. Single-cell V(D)J sequencing captures this diversity one cell at a time, preserving the native pairing of receptor chains that bulk repertoire sequencing irreversibly scrambles.

The 10x Genomics Chromium Single Cell Immune Profiling workflow, using the 5' v2 or v3 chemistry, enriches full-length V(D)J transcripts via target enrichment primers that bind to the constant regions of TCR and BCR transcripts. Within each GEM (gel bead in emulsion), mRNA and V(D)J transcripts from the same cell are captured by the same gel bead, sharing a common 10x cell barcode. After reverse transcription, the V(D)J library is enriched by PCR using primers specific to the constant regions of TCR α/β or BCR heavy/light chains. The resulting library yields paired, full-length receptor sequences — both chains from the same cell, preserving the native α-β or heavy-light pairing that determines antigen specificity. Sequencing at 5,000 read pairs per cell is typically sufficient for confident clonotype assembly, though deeper sequencing (10,000 to 25,000 reads per cell) improves detection of rare clonotypes and somatic hypermutation calling in BCR sequences.

The key metrics derived from a single-cell V(D)J dataset include clonotype frequency (what fraction of cells express each unique receptor), clonal expansion (which clonotypes are enriched above naive baseline), repertoire diversity (Shannon entropy or inverse Simpson index measuring the evenness of clonotype distribution), and somatic hypermutation load (for BCRs, the number of nucleotide substitutions relative to germline, which correlates with affinity maturation). For T cells, the complementarity-determining region 3 (CDR3) — the most variable portion of the receptor at the V(D)J junction — is the primary determinant of antigen recognition, and CDR3 length distribution and amino acid composition are standard repertoire-level metrics used to compare immune states across conditions.

TCR and BCR Clonotype AssemblyFigure 2: TCR and BCR Clonotype Assembly — Paired Chain Capture, CDR3 Identification, and Clonal Expansion Detection

CITE-Seq — Surface Proteomics Meets Transcriptomics

Transcriptional profiles alone cannot fully resolve immune cell identity. A CD8+ effector memory T cell and a CD8+ exhausted T cell share transcriptional programs but differ in surface protein expression in ways that are functionally decisive — PD-1, TIM-3, and LAG-3 protein levels define exhaustion better than their transcripts, which are subject to post-transcriptional regulation. CITE-Seq addresses this gap by using oligonucleotide-conjugated antibodies — commercialized as TotalSeq by BioLegend — that carry a DNA barcode instead of a fluorophore. Cells are stained with a cocktail of TotalSeq antibodies before encapsulation in the 10x Chromium chip. Within each GEM, both the polyadenylated mRNA and the antibody-derived tag (ADT) oligos are captured and reverse-transcribed, producing two distinct libraries from the same single cells: a gene expression (GEX) library and an ADT library.

The practical advantage is that surface protein detection via ADTs is not constrained by the same technical limitations as transcript-based immunophenotyping. Lowly expressed transcripts such as those encoding CD25 (IL2RA) or CD127 (IL7R) may produce only a handful of UMIs per cell — near the stochastic detection limit — while the corresponding ADT signal from surface protein is often 10- to 100-fold higher in signal-to-noise ratio. Pre-titrated TotalSeq panels covering 30 to 200 markers are now standard, including the widely adopted TotalSeq Human TBNK panel (CD3, CD4, CD8, CD11c, CD14, CD16, CD19, CD45, CD56) and the TotalSeq Universal Cocktail. For projects requiring 10x Genomics-based single-cell immune profiling, the standard approach is to combine a TotalSeq-C panel (compatible with 5' chemistry) with the V(D)J enrichment workflow, yielding three data layers — transcriptome, immune repertoire, and surface proteome — from every cell.

Cell Hashing is a complementary application of the same oligo-conjugated antibody technology. By staining each sample with a unique hashtag oligonucleotide (HTO) antibody targeting ubiquitously expressed surface markers such as CD298 or β2-microglobulin, multiple samples can be pooled into a single 10x lane and demultiplexed computationally. Current protocols support multiplexing of up to 12 to 14 hashtag-labeled samples in a single 10x lane with balanced recovery, reducing per-sample library preparation cost and eliminating batch effects between samples run in the same lane — a critical consideration for longitudinal vaccine studies or multi-timepoint immunotherapy monitoring. For multi-sample immune profiling studies, single-cell data analysis that integrates HTO demultiplexing with clonotype assembly and ADT normalization ensures that each cell's repertoire and surface proteome are correctly assigned to its sample of origin before downstream biological interpretation.

CITE-Seq Feature Barcoding MechanismFigure 3: CITE-Seq Feature Barcoding Mechanism — TotalSeq Antibody Staining, GEM Capture, and ADT Library Construction

BEAM — Adding Antigen Specificity to the Multimodal Picture

The most significant recent advance in single-cell immune profiling is BEAM (Barcode Enabled Antigen Mapping), which adds a fourth layer — antigen specificity — to the existing V(D)J + GEX + ADT stack. BEAM uses Feature Barcode-conjugated antigens: for T cells (BEAM-T), fluorescently labeled peptide-MHC monomers loaded with the peptide of interest are conjugated to DNA-barcoded dextramers; for B cells (BEAM-Ab), the antigen protein is directly conjugated. Antigen-specific cells are enriched by flow sorting before encapsulation, and the BEAM barcode is captured alongside the V(D)J, GEX, and ADT libraries within the same GEM.

The Cell Ranger multi pipeline, available from version 7.1 and refined through version 8.0 in 2025, processes all four library types simultaneously. The antigen specificity algorithm uses a beta-distribution model to score each cell barcode: (1 − beta.cdf(0.925, Antigen_UMI + 1, Control_UMI + 3)) × 100, producing a score from approximately 0 to 100 that separates antigen-binding cells from background. The score does not measure binding affinity — it assesses whether the cell's receptor associates with the antigen above the negative control background — but when combined with paired TCR/BCR sequences, it directly links a clonotype to its cognate antigen. This capability is particularly powerful in immuno-oncology, where identifying the specific TCR sequences that recognize a tumor neoantigen enables engineered TCR-T cell therapy development.

The current limitations of BEAM are worth noting. It is incompatible with GEM-X v3 chemistry and requires the legacy 5' v2 assay. The 16 available BEAM conjugate barcodes constrain multiplexed antigen panels. And because antigen-specific cells must be enriched by flow sorting, the input cell number and sorting strategy must be optimized for each antigen of interest — a practical consideration that affects experimental design at the outset. For investigators who need antigen specificity mapping without building the full BEAM workflow in-house, 10x single-cell immune profiling services that include BEAM sample processing, flow sorting enrichment, and antigen specificity scoring provide an end-to-end pipeline from sorted cells to annotated clonotype-antigen pairs.

Analysis Tools for Multimodal Immune Profiling

The computational ecosystem for integrated V(D)J and CITE-Seq analysis has matured considerably in the 2025-2026 period. The Cell Ranger multi pipeline is the starting point, handling alignment (STAR for GEX, a customized assembly algorithm for V(D)J), clonotype grouping, antigen specificity scoring, and ADT count matrix generation. Loupe V(D)J Browser v5.2 supports interactive exploration of clonotypes colored by antigen specificity scores, chain-pairing visualization, and aggregated analysis across multiple BEAM libraries — making it possible to compare antigen-specific clonotypes across timepoints or tissue compartments without opening a command line.

For programmatic analysis, the R package scRepertoire provides a Seurat-compatible framework for clonotype visualization, including clonal expansion tracking, clonal overlap between conditions, and diversity metric calculation. The package supports loading Cell Ranger V(D)J outputs directly and integrates with the standard Seurat object through barcode matching — clonotype data is stored as metadata columns, and functions such as clonalQuant, clonalDiversity, and clonalOverlap produce publication-ready visualizations with minimal code. The Python package Dandelion (v0.5.7, January 2026) integrates BCR/TCR analysis with Scanpy and supports trajectory-based clonotype analysis, somatic hypermutation calling, and isotype-switch reconstruction from BCR sequences. MiXCR v4.7 supports single-cell V(D)J assembly and repertoire analysis from 10x Chromium data with dedicated presets for paired-chain immune profiling, covering the full workflow — V(D)J assembly, ADT count normalization, and joint clustering — from a single command-line invocation. For researchers requiring custom bioinformatics support for immune profiling, the ability to integrate clonotype, transcriptome, and surface protein data within a unified analytical framework is essential for extracting biological meaning from the raw count matrices.

Data integration across modalities is typically performed using weighted nearest neighbor (WNN) analysis in Seurat v5, which learns cell-type-specific modality weights from the GEX and ADT modalities and constructs a joint neighbor graph. For V(D)J integration, clonotype information is typically overlaid as metadata — coloring UMAP clusters by clonotype frequency, highlighting expanded clonotypes, or subsetting to antigen-specific cells identified by BEAM. The critical QC consideration in multimodal immune profiling is that each modality carries its own failure modes: V(D)J libraries can drop one chain (producing unpaired receptor sequences), ADT libraries can suffer from ambient antibody signal, and BEAM libraries can produce false-positive antigen calls from nonspecific dextramer binding. Cross-modality QC — for example, confirming that a cell identified as antigen-specific by BEAM also shows an activated transcriptional signature and surface protein markers consistent with recent TCR engagement — increases confidence that the multimodal annotation reflects genuine biology rather than technical artifact.

Multimodal Immune Profiling Analysis PipelineFigure 4: Multimodal Immune Profiling Analysis Pipeline — From Cell Ranger multi to Integrated Loupe/Seurat Visualization

Applications in Immunology and Drug Development

The strongest application of combined V(D)J and CITE-Seq is in immuno-oncology, where the technology simultaneously answers three questions about every T cell in a tumor: what receptor does it express (clonotype), what proteins are on its surface (phenotype), and what does it recognize (antigen specificity with BEAM). A 2025 study profiling tumor-infiltrating lymphocytes from melanoma patients identified neoantigen-specific CD8+ T cell clonotypes that co-expressed PD-1, TIM-3, CD39, and CD103 at the protein level — a surface phenotype consistent with tumor-reactive, tissue-resident memory cells — and tracked their clonal expansion following checkpoint blockade. Without CITE-Seq, the same study would have relied on transcript-based inference of protein expression, which correlates only moderately (Spearman ρ ≈ 0.5-0.7) with actual surface protein abundance for key immune checkpoint molecules. Integrating these multimodal immune profiling approaches with single-cell ATAC-seq adds chromatin accessibility profiling, enabling a regulatory-dimension view of the same clonotypes — linking the TCR sequence to the transcription factor binding events that drive exhaustion or effector programs.

Vaccine research is the second major application domain. Longitudinal BCR repertoire profiling with paired surface proteomics can track antigen-specific B cell clonal expansion, isotype switching, and somatic hypermutation accumulation following immunization. CITE-Seq surface markers — CD27, CD38, CD138, IgD — resolve plasma cell, memory B cell, and naive B cell populations with higher resolution than transcript-based clustering alone, enabling precise isolation of vaccine-responding populations. In a SARS-CoV-2 mRNA vaccine study design, combining V(D)J and CITE-Seq with Cell Hashing reduced the per-sample cost by multiplexing 12 timepoints from three donors across two 10x lanes, while eliminating inter-lane batch effects that would otherwise confound the comparison of pre-boost and post-boost timepoints.

Autoimmune disease research and therapeutic antibody discovery represent emerging applications. In systemic lupus erythematosus, paired BCR sequencing and surface phenotyping can identify autoreactive B cell clonotypes based on their V-gene usage and somatic hypermutation patterns, then link them to surface protein signatures that distinguish pathogenic from bystander populations. For therapeutic antibody discovery, combining BCR sequencing with CITE-Seq enables screening of antigen-specific B cells — identified by BEAM-Ab — for surface markers associated with high-affinity antibody secretion (CD27^hi, CD38^hi, CD138+), prioritizing clonotypes for recombinant expression and functional validation.

An emerging recognition in the field is that the primary bottleneck in multimodal immune profiling has shifted from data generation to data integration. Each modality — transcriptome, V(D)J, ADT, and BEAM — carries distinct noise structures and batch effects that do not align across modalities. When 12 to 14 Cell-Hashed samples are pooled and demultiplexed, the HTO demultiplexing errors propagate into all downstream modalities, meaning a misassigned cell in the transcriptome is also misassigned in the clonotype and surface protein tables. Tools such as Seurat v5 WNN and the multi-omics factor analysis (MOFA) framework now support joint modeling of three or more modalities, learning latent factors that capture coordinated variation across transcriptome, surface proteome, and clonotype space. For BEAM data in particular, integrating antigen specificity scores with the paired TCR sequence and surface activation markers into a unified latent representation remains an active area of methods development, with the potential to link receptor sequence, protein phenotype, and antigen target into a single predictive model of T cell function.

For investigators designing multimodal immune studies, a practical recommendation is to define which modalities are essential for the primary endpoint and which are exploratory. A study whose primary readout is clonal expansion after therapy needs V(D)J + GEX; adding CITE-Seq enables deeper phenotypic characterization but increases per-cell sequencing cost by 20 to 30 percent. A study whose primary readout is surface protein-based cell-type quantification across conditions needs CITE-Seq + GEX; adding V(D)J is valuable only if the biological question involves clonotype tracking. BEAM should be reserved for studies where antigen specificity is the central biological question, given its requirement for the legacy 5' v2 chemistry, custom BEAM conjugate preparation, and flow-sorting enrichment. Sample input for BEAM experiments should be planned conservatively: antigen-specific precursor frequencies are often below 0.1 percent in unexpanded T cell populations, meaning 5 to 10 million total input cells may be needed to recover a few hundred antigen-specific cells after sorting — a consideration that directly affects tissue collection and sample allocation. For a broader view of how single-cell and spatial technologies fit into a unified experimental strategy, see our hub guide on Single-Cell and Spatial Biology Services.

V(D)J and CITE-Seq ApplicationsFigure 5: Applications of V(D)J and CITE-Seq — Immuno-Oncology, Vaccine Research, and Therapeutic Antibody Discovery

FAQ

What is the difference between 5' and 3' single-cell chemistry for immune profiling?

The 5' chemistry captures transcripts from the 5' end and includes V(D)J enrichment primers targeting the constant regions of TCR and BCR transcripts, making it the required chemistry for immune repertoire profiling. The 3' chemistry captures from the poly(A) tail and does not support V(D)J enrichment, but provides higher transcriptomic sensitivity for gene expression analysis. For combined V(D)J + CITE-Seq, the 5' v2 chemistry with TotalSeq-C antibodies is the standard workflow.

How many surface proteins can I profile with CITE-Seq?

Current TotalSeq panels support up to 200 markers simultaneously, though most studies use 30 to 50 antibodies focused on the immune lineages and functional markers relevant to their biological question. The practical limit is set by antibody availability, spectral overlap-free barcode design, and the diminishing returns of adding markers that do not improve cell-type discrimination beyond what the core panel already provides.

What is Cell Hashing and when should I use it?

Cell Hashing uses oligonucleotide-conjugated antibodies against ubiquitously expressed surface markers (CD298, β2-microglobulin) to label each sample with a unique barcode before pooling. It enables multiplexing of up to 14 samples in a single 10x lane, reducing per-sample cost and eliminating batch effects. Use it for multi-sample, multi-condition, or multi-timepoint studies where batch correction would otherwise complicate the analysis.

Can I do V(D)J sequencing on non-human/non-mouse species?

The 10x V(D)J enrichment primers are designed for human and mouse constant regions. Cross-reactivity has been reported for some non-human primate and rat samples, but most non-model species require custom primer design targeting the constant regions of their TCR/BCR loci. For any non-standard species, consult with the service provider about custom primer feasibility before committing samples.

What is the minimum number of cells needed for V(D)J sequencing?

A minimum of 5,000 to 10,000 viable T or B cells is recommended for meaningful clonotype diversity analysis, though clonal expansion — where a few dominant clonotypes occupy a large fraction of the repertoire — can be detected with as few as 1,000 cells. For rare populations such as antigen-specific T cells (pre-enrichment), higher input numbers (50,000 to 100,000 total cells) are recommended to ensure adequate representation of the population of interest.

How does CITE-Seq compare to flow cytometry for surface protein detection?

CITE-Seq detects surface proteins via DNA barcode readout rather than fluorescence, enabling simultaneous measurement of 30 to 200 markers — far exceeding the roughly 18 to 30 parameters achievable with spectral flow cytometry. However, CITE-Seq destroys the cell and cannot sort live populations. The two technologies are complementary: flow cytometry for live-cell isolation and sorting, CITE-Seq for high-parameter discovery and integration with transcriptomic and repertoire data.

What sequencing depth is recommended for combined V(D)J + CITE-Seq?

Gene expression library: 20,000 to 25,000 read pairs per cell. V(D)J library: 5,000 to 10,000 read pairs per cell. ADT library: 5,000 to 10,000 read pairs per cell. For BEAM-Antigen Capture libraries, 5,000 reads per cell. At 10,000 cells per sample, total sequencing is approximately 350 to 500 million read pairs.

What is the turnaround time for a V(D)J + CITE-Seq project?

Sample processing through library construction typically requires 2 to 3 weeks, and sequencing 1 to 2 additional weeks. Full bioinformatics analysis — including clonotype assembly, ADT normalization, multimodal clustering, and antigen specificity scoring — adds 2 to 4 weeks.

References:

  1. Stoeckius M, Hafemeister C, Stephenson W, et al. Simultaneous epitope and transcriptome measurement in single cells. Nature Methods. 2017;14(9):865-868. https://doi.org/10.1038/nmeth.4380
  2. Hao Y, et al. Integrated analysis of multimodal single-cell data. Cell. 2021;184(13):3573-3587. https://doi.org/10.1016/j.cell.2021.04.048
  3. Mimitou EP, Cheng A, Montalbano A, et al. Multiplexed detection of proteins, transcriptomes, clonotypes and CRISPR perturbations in single cells. Nature Methods. 2019;16(5):409-412. https://doi.org/10.1038/s41592-019-0392-0
  4. Yost KE, Satpathy AT, Wells DK, et al. Clonal replacement of tumor-specific T cells following PD-1 blockade. Nature Medicine. 2019;25(8):1251-1259. https://doi.org/10.1038/s41591-019-0522-3
  5. Zheng GXY, Terry JM, Belgrader P, et al. Massively parallel digital transcriptional profiling of single cells. Nature Communications. 2017;8:14049. https://doi.org/10.1038/ncomms14049
  6. Stubbington MJT, Lonnberg T, Proserpio V, et al. T cell fate and clonality inference from single-cell transcriptomes. Nature Methods. 2016;13(4):329-332. https://doi.org/10.1038/nmeth.3800
  7. Song HW, Martin J, Shi X, Tyznik AJ. Key considerations on CITE-Seq for single-cell multiomics. Proteomics. 2025;25(21-22):206-213. https://doi.org/10.1002/pmic.202400011

For research purposes only, not intended for clinical diagnosis, treatment, or individual health assessments.

For research purposes only, not intended for clinical diagnosis, treatment, or individual health assessments.
Speak to Our Scientists
What would you like to discuss?
With whom will we be speaking?

* is a required item.

Contact CD Genomics
Terms & Conditions | Privacy Policy | Feedback   Copyright © CD Genomics. All rights reserved.
Top