
Antibody discovery depends on recovering the sequence relationships that connect B-cell diversity to a usable candidate. Short-read repertoire sequencing can profile large numbers of rearrangements, but fragmented reads may separate V(D)J sequence from constant-region context, obscure long antibody constructs, or require reconstruction when the project needs a complete sequence rather than a short clonotype tag.
CD Genomics provides long-read sequencing for antibody discovery and development using complementary PacBio HiFi and Oxford Nanopore Technologies (ONT) workflows. We support full-length BCR repertoire profiling, clonotype and somatic-hypermutation analysis, single-cell strategies for native heavy-light pairing, long antibody-construct sequencing, and sequencing of display-library outputs for candidate tracking and prioritization.
Solution highlights
The sequence diversity that makes antibodies useful also makes them difficult to reconstruct. V(D)J recombination, junctional diversity, somatic hypermutation, class-switch recombination, and alternative transcript structures create many related molecules that may differ by only a small number of nucleotides. The discovery question is therefore rarely just “which CDR3 is present?” It may be “which complete variable region belongs to this lineage?”, “which isotype carries the clone?”, “which heavy and light chains came from the same cell?”, or “which full-length display construct was enriched after selection?”
Long-read sequencing is valuable when the linkage between sequence features matters. A single read can cover the complete V(D)J region and extend into constant-region sequence, enabling isotype or subclass assignment without stitching together separate fragments. In display libraries, one read can span an entire scFv or other antibody insert, preserving VH, linker, and VL order on one molecule. In single-cell workflows, long reads can retain upstream cell barcodes and full-length transcript sequence so that candidate sequences remain connected to cellular context.
Our Full-Length TCR/BCR Repertoire Profiling service is the core targeted immune-repertoire entry point. This solution page extends that capability into an antibody-development workflow by showing when bulk repertoire sequencing, single-cell pairing, full-length transcript sequencing, or display-library sequencing is the appropriate evidence layer.
Long reads are not a universal replacement for short-read immune repertoire sequencing. Short reads can be more economical when the objective is very deep counting of a predefined short V(D)J amplicon across large cohorts. We recommend long reads when sequence completeness, isotype linkage, lineage resolution, chain pairing, antibody-construct integrity, or difficult-to-assemble repertoire regions are central to the decision.
| Research question | Long-read evidence | Decision supported |
| Which B-cell clones expanded after immunization or antigen exposure? | Full-length V(D)J sequences, clonotype counts, SHM profiles, isotype/subclass context | Prioritize expanded and affinity-matured lineages for follow-up. |
| How did an antibody lineage evolve? | Complete variable-region sequences across samples or time points | Reconstruct lineage trees and identify sequence variants accumulated during maturation. |
| Which heavy and light chains belong together? | Single-cell or barcode-preserving full-length transcript data | Recover native VH–VL pairs for recombinant expression and functional testing. |
| Which antibody-library inserts are enriched during selection? | Full-length scFv, Fab, VHH, or other library-insert reads | Track enrichment, diversity loss, sequence families, and low-frequency candidate persistence. |
| Is a candidate antibody sequence complete and unambiguous? | Full-length cDNA, amplicon, or construct-spanning reads | Confirm variable-region sequence, chain identity, ORF integrity, and construct architecture. |
| Does germline IG variation complicate repertoire interpretation? | Optional long-read genomic analysis of complex IG loci | Improve allele assignment and distinguish true SHM from germline variation when required. |
| Which candidates should move to functional testing? | Integrated clonotype abundance, SHM, lineage, pairing, isotype, and sequence-quality evidence | Build a transparent sequence-based candidate shortlist for orthogonal validation. |
The output is a sequence-based evidence package. Binding affinity, neutralization, epitope specificity, expression yield, aggregation, immunogenicity, and manufacturability cannot be concluded from repertoire sequencing alone and require orthogonal functional or biophysical assays.

Bulk BCR repertoire sequencing is a high-depth way to interrogate the antibody response in PBMCs, B-cell-enriched fractions, blood, bone marrow, lymphoid tissues, or other research samples containing B cells. Long reads can cover complete immunoglobulin variable domains and extend into constant-region sequence, allowing a clone to be interpreted with its V, D, J, CDR1, CDR2, CDR3, framework, and isotype information connected on the same molecule.
We annotate V(D)J gene usage, CDR3 sequence, productive versus nonproductive rearrangements, and clonal clusters. Comparing pre- and post-immunization samples, tissues, treatment groups, or longitudinal time points can identify lineages that expand, contract, persist, or disseminate between compartments. The goal is not to assume that the most abundant clone is automatically the best antibody, but to identify sequence families with evidence that justifies functional follow-up.
Antigen-experienced B cells accumulate somatic mutations as lineages evolve. High-accuracy long reads are useful when closely related sequence variants must be separated without confusing sequencing error for biological mutation. Lineage reconstruction can group related VH sequences, estimate distance from inferred germline, and show how mutations accumulate across samples or time. For high-confidence SHM analysis, we favor accuracy-aware designs and careful germline assignment rather than treating every mismatch as a true biological event.
Full-length reads can retain constant-region sequence adjacent to the antigen-binding variable region. This allows an expanded clonotype to be interpreted together with IgM, IgD, IgG, IgA, or other available constant-region context, depending on species, library design, transcript structure, and read length. Class switching can therefore become part of lineage interpretation rather than a separate assay.
For projects that need broader transcript context around B cells rather than targeted repertoire depth alone, Full-Length Transcript Sequencing (Iso-Seq) can provide isoform-resolved RNA information, while a targeted BCR assay remains the more efficient route when immune-receptor sequence depth is the primary objective.
A critical distinction in antibody discovery is that full-length BCR sequencing does not automatically mean native heavy-light pairing. Bulk repertoire libraries usually sequence heavy and light chains as independent molecules. They can reveal deep clonal structure, but a VH sequence and a VL sequence observed in the same bulk sample cannot be assumed to originate from the same B cell.
When native pairing is required, the project must preserve cell identity or molecular linkage before long-read sequencing. Our Single-Cell Full-Length Transcriptome Sequencing workflows provide a route for retaining cell barcodes while resolving complete transcripts. With an appropriate upstream V(D)J or full-length cDNA design, heavy- and light-chain sequences can be assigned back to the same cell and linked to cell state, expression phenotype, or other transcriptomic information.
| Design | Primary strength | Main limitation | Best use |
| Bulk full-length BCR | Deep repertoire sampling and lineage analysis | Native VH–VL pairing generally not preserved | Clonotype expansion, SHM, isotype, longitudinal lineage tracking |
| Single-cell paired-chain | Native heavy-light pairing with cell identity | Lower repertoire depth and higher per-cell complexity than bulk | Candidate reconstruction, antigen-specific B-cell studies, paired-chain discovery |
| Sorted single/few-cell targeted recovery | Direct pairing from selected B cells | Lower throughput | Focused recovery of candidates from phenotypically selected cells |
| Linked antibody construct library | VH–VL relationship physically encoded in one scFv/Fab/VHH construct | Represents the library design, not necessarily native B-cell pairing | Display-library sequencing and candidate enrichment tracking |
Pairing strategy should be decided before library preparation. If the project begins with already-generated bulk BCR reads, no downstream algorithm can reliably recreate native heavy-light pairing that was never encoded in the library.
Antibody development frequently generates linked constructs longer than a conventional short-read amplicon. scFv libraries, Fab-related constructs, VHH libraries, synthetic antibody pools, and engineered variants may contain the complete antigen-binding sequence plus linkers, framework changes, barcodes, or vector-adjacent sequence. Long reads can span the full insert so that the candidate is observed as one sequence rather than computationally reconstructed from separate fragments.
If you provide DNA or cDNA from pre-selection libraries and subsequent phage, yeast, or other display outputs, sequencing can compare clone frequencies and sequence families across rounds. This supports enrichment tracking, identification of convergent sequence motifs, detection of diversity bottlenecks, and recovery of candidates that remain below the most abundant clones. A 2024 independent study demonstrated the use of high-accuracy ONT sequencing with dual molecular barcodes to monitor antibody phage-display diversity and enrichment and to recover rare binders that conventional colony picking could miss.
For a selected candidate, a complete long read can verify the order and sequence of variable domains, linker regions, framework sequence, and other encoded elements captured by the assay. This is useful before synthesis or expression when the library contains closely related variants or when the selected sequence is too long to cover confidently with one conventional short read.
CD Genomics provides the sequencing and sequence-analysis layer. We do not infer antigen binding, affinity, specificity, or developability from sequence enrichment alone. A candidate that rises strongly through panning still requires appropriate recombinant expression and functional validation.
| Decision factor | PacBio HiFi | Oxford Nanopore Technologies |
| Priority | High per-read consensus accuracy for closely related antibody sequences | Flexible read length, rapid targeted sequencing, and long construct coverage |
| Strong antibody use case | Full-length BCR repertoire, SHM analysis, clonotype discrimination, accurate candidate sequence recovery | Display-library insert sequencing, rapid construct review, custom long amplicons, flexible library formats |
| Near-identical clone resolution | Well suited when single-nucleotide differences matter | Requires accuracy-aware basecalling, consensus, molecular barcodes, or replicate strategy when single-nucleotide calls drive selection |
| Very long or heterogeneous inserts | Strong for HiFi-compatible insert ranges | Particularly flexible for long and heterogeneous molecules |
| Turnkey targeted repertoire context | Natural fit with full-length immune-repertoire and Iso-Seq workflows | Useful where rapid long-read amplicon or construct sequencing is the primary requirement |
Our PacBio SMRT Sequencing Technology and Oxford Nanopore Sequencing Technology pages describe the broader platform characteristics. For antibody projects, the choice is made at the level of the biological decision: accuracy for SHM and closely related clonotypes, pairing design, insert length, required depth, and whether the data come from a natural B-cell repertoire or an engineered library.
We begin with the candidate decision rather than the sequencer. The same project may use deep bulk repertoire sequencing to identify lineages and a smaller single-cell experiment to recover paired candidates, or long-read sequencing of a display library followed by focused verification of selected clones.

We clarify whether the project needs repertoire depth, native heavy-light pairing, lineage evolution, isotype context, display-library enrichment, full construct sequence, or a combination of these outputs.
Inputs may include B-cell-rich biological material, purified RNA or cDNA, pre-generated single-cell libraries, targeted BCR amplicons, display-library DNA/cDNA, or candidate constructs. The library design is chosen to preserve the linkage information that the downstream decision requires.
PacBio HiFi or ONT sequencing is selected based on required accuracy, insert length, throughput, and project format. Controls and technical replicates can be incorporated when low-frequency candidates or sequence-level error discrimination is important.
Reads are processed for V(D)J assignment, CDR annotation, productive status, isotype or subclass, SHM, clonotype clustering, heavy-light pairing when encoded, or complete display-construct parsing.
We compare abundance, sequence similarity, SHM, lineage structure, sample distribution, selection-round enrichment, and other relevant evidence to identify candidates or sequence families that merit follow-up.
Final outputs include candidate nucleotide and translated amino-acid sequences, lineage or pairing evidence, sequence QC, annotations, visualizations, and a clear statement of which conclusions require downstream functional validation.
| Analysis module | Typical outputs | Discovery value |
| Read QC and consensus processing | Length, quality, full-length status, molecular barcodes/consensus metrics where applicable | Separates sequence evidence from platform or PCR noise. |
| V(D)J and CDR annotation | IGHV/IGHD/IGHJ or light-chain assignments, CDR1/2/3, framework regions | Defines the complete antigen-binding variable sequence. |
| Productive sequence review | ORF status, stop codons, frameshifts, complete variable-domain sequence | Prevents nonproductive or damaged sequences from entering candidate lists. |
| Clonotype and diversity analysis | Clone frequencies, richness/diversity, V/J usage, CDR3 distributions | Identifies expansion and repertoire shifts. |
| Somatic hypermutation | Mutation counts, germline distance, mutation maps, lineage-specific changes | Supports affinity-maturation and lineage interpretation. |
| Lineage reconstruction | Sequence clusters, phylogenetic trees, temporal/tissue distribution | Finds related candidate variants and maturation paths. |
| Isotype/subclass analysis | Constant-region assignment linked to full variable sequence | Adds class-switch context to antigen-binding clones. |
| Heavy-light pairing | Per-cell or molecularly linked VH–VL pairs when the library preserves pairing | Produces reconstructable native antibody candidate sequences. |
| Display-library mining | Full insert sequences, enrichment trajectories, family clusters, rare-candidate persistence | Supports sequence-driven clone selection across panning or selection rounds. |
| Candidate shortlist | Ranked sequence table with evidence fields and QC flags | Provides transparent inputs for downstream functional screening. |
Where germline ambiguity is a major source of uncertainty, optional genomic long-read analysis can resolve complex IG loci and improve the repertoire reference. A large Nature Communications study of 154 individuals showed that long-read IGH genotyping uncovered extensive SNV, indel, and structural variation and demonstrated that IGH genotype can strongly shape expressed antibody gene usage. This is particularly relevant when apparent “mutation” or unusual gene usage may reflect germline diversity rather than antigen-driven evolution.

| Entry mode | Common material | Best suited for | Planning note |
| Bulk BCR repertoire | PBMCs, B-cell-enriched fractions, whole blood, bone marrow, fresh/frozen tissue, RNA or cDNA | Deep clonotype, SHM, isotype, and lineage profiling | Native heavy-light pairing is generally not preserved in bulk libraries. |
| Single-cell paired-chain | Viable cell suspensions or compatible pre-generated single-cell cDNA/library material | Native VH–VL pairing and cell-state linkage | Upstream barcode compatibility must be reviewed before sequencing. |
| Sorted single/few B cells | Phenotypically or antigen-selected B cells | Focused paired candidate recovery | Lower throughput but direct connection to selected cell phenotype. |
| Display-library sequencing | Phage/yeast/other display plasmid DNA, PCR products, or cDNA from selection rounds | Full insert sequencing and enrichment tracking | Provide library architecture, expected insert range, round labels, and barcode information. |
| Candidate construct verification | Purified plasmid DNA, cDNA, or full-length amplicon | Sequence confirmation before downstream expression | Reference construct sequence and expected antibody format improve interpretation. |
Exact input quantity, RNA integrity, cell viability, library concentration, and sequencing depth depend on the selected workflow and sample type. We confirm acceptance criteria after reviewing the project design rather than applying one generic threshold to all antibody-discovery projects.
For custom full-length cDNA projects, Nanopore Full-Length cDNA Sequencing can provide an ONT-based entry route when flexible transcript length and custom cDNA analysis are priorities.
We support both PacBio and ONT strategies rather than forcing every antibody project into one sequencing format. This allows us to prioritize HiFi accuracy for closely related repertoire sequences or choose flexible ONT workflows for long constructs and custom library formats.
The LongSeq service portfolio connects targeted full-length BCR profiling with single-cell full-length transcriptomics and broader long-read RNA workflows. This makes it possible to design a tiered study in which deep bulk sequencing identifies lineages and a focused single-cell arm recovers native paired candidates.
Our analysis does more than count reads. Candidate interpretation can integrate V(D)J assignment, CDR sequence, productive status, isotype, SHM, lineage structure, pairing evidence, tissue or time-point distribution, and display-round enrichment while keeping technical uncertainty visible.
We do not label an antibody “high affinity,” “neutralizing,” “specific,” “developable,” or “manufacturable” from sequencing alone. Instead, we deliver the sequence-defined candidate set and the evidence used to prioritize it, so that downstream expression and functional assays can be designed around traceable sequence information.
Crescioli S, Correa I, Ng J, et al. B cell profiles, antibody repertoire and reactivity reveal dysregulated responses with autoimmune features in melanoma. Nature Communications. 2023;14:3378. doi:10.1038/s41467-023-39042-y.
The investigators sought to characterize tumor-resident and circulating B-cell responses in melanoma, including antibody repertoire structure and the reactivity of patient-derived antibodies. The study combined immune phenotyping, antibody sequence analysis, PacBio long-read repertoire sequencing, recombinant antibody production, and antigen-reactivity experiments.
For high-throughput tissue antibody repertoires, the authors generated Ig-targeted long-read libraries and sequenced them using PacBio Sequel II. Full-length antibody sequences were annotated for V(D)J features, heavy-chain subclass, CDR3 properties, clonal structure, and lineage relationships. In parallel, matched heavy- and light-chain variable regions recovered from single sorted tumor-resident memory B cells were cloned into IgG1 expression vectors for downstream recombinant antibody testing.
Long-read repertoire analysis showed distinct V, D, and J usage and isotype-associated patterns between matched blood and tumor samples. The authors also reported more unproductive heavy- and light-chain sequences in tumor than blood in the matched long-read repertoire analysis, supporting a distinct tumor-associated B-cell compartment. Downstream, twelve patient-derived recombinant IgG1 antibodies were produced from matched heavy- and light-chain sequences. Six of the twelve pulled down possible autoantigens from both skin and melanoma tissue lysates, and two showed reactivity to carbohydrate targets in glycan-array testing.
Published Figure 4 illustrates matched blood-versus-tumor antibody repertoire features, including isotype distribution, V(D)J usage, and variable-domain clustering.
This independent example demonstrates an important antibody-discovery principle: long-read repertoire sequencing can define complete sequence families and clonal context, while the biological meaning of those sequences is established by downstream recombinant expression and functional testing. The sequencing layer narrows and organizes the candidate space; it does not replace antigen-binding or specificity experiments.
It can recover complete heavy- or light-chain variable sequences and often their constant-region context, but a complete native antibody candidate requires correct VH–VL pairing. Bulk repertoire sequencing does not generally preserve that pairing. Use a single-cell, barcoded, or physically linked library design when native heavy-light reconstruction is required.
PacBio HiFi is particularly useful when accurate discrimination of closely related clonotypes and somatic-hypermutation variants is important. Consensus accuracy helps reduce the risk that sequencing error is mistaken for a biologically meaningful mutation.
ONT is useful for flexible long-amplicon and antibody-construct sequencing, including display-library inserts. For decisions that depend on single-nucleotide differences, the study should use an accuracy-aware strategy such as high-accuracy basecalling, consensus, molecular barcodes, or orthogonal confirmation rather than treating every raw-read difference as a true variant.
Yes, as a sequencing and analysis project when you provide the library DNA, cDNA, or full-length amplicons. Long reads can span complete linked antibody inserts and compare sequence enrichment across selection rounds. The display selection, binding assay, and functional validation are separate activities and are not inferred from sequencing alone.
No. Expansion, somatic hypermutation, lineage persistence, or selection-round enrichment can prioritize candidates, but affinity, specificity, neutralization, epitope, potency, and developability require orthogonal functional and biophysical measurements.
Yes. A practical design is to use bulk full-length BCR sequencing for deep lineage discovery and longitudinal tracking, then use a focused single-cell or antigen-selected arm to recover native heavy-light pairs from the most relevant populations. The two datasets answer complementary questions.
Selected model species can be supported after review of germline references, primer strategy, and expected immunoglobulin architecture. Non-human repertoire projects require more care because germline databases may be incomplete; long reads can help improve sequence context, but reference limitations must be included in interpretation.
1. Full-length antibody sequence table — productive heavy/light sequences with V(D)J, CDR1/2/3, framework, isotype/subclass, nucleotide, and translated amino-acid annotations.
2. Clonotype and lineage report — clone abundance, diversity, SHM, lineage trees, sample/tissue/time-point distribution, and candidate sequence families.
3. Paired-chain candidate table — VH–VL pairs with cell/barcode provenance when the experimental design preserves native pairing.
4. Display-library enrichment report — full insert sequences, counts by selection round, enrichment trajectories, sequence-family clustering, and rare-candidate persistence.
5. Candidate prioritization matrix — sequence-based evidence used to shortlist clones for downstream recombinant expression and functional testing, with QC flags and limitations stated explicitly.
References
For Research Use Only. Not for use in diagnostic or clinical procedures.