Long-Read Sequencing for Gene Therapy and AAV Characterization

Long-Read Sequencing for Gene Therapy and AAV Characterization

long-read sequencing solution for AAV vector genome and gene therapy characterization

Gene-therapy vector characterization becomes difficult when the question is not simply whether a transgene sequence is present, but whether the complete vector genome is packaged as designed, which partial or rearranged forms coexist in the preparation, whether plasmid or host-cell sequences are present, and how vector DNA is organized after delivery. These are fundamentally long-range sequence questions.

CD Genomics provides long-read sequencing for gene therapy and AAV characterization using complementary PacBio HiFi and Oxford Nanopore Technologies (ONT) workflows. We help research and biopharma teams characterize packaged recombinant AAV (rAAV) genomes, plasmid templates, complex vector forms, host-vector junctions, integration architecture, and selected epigenetic features with read-level structural evidence that can span entire vector molecules or long host-genome contexts.

Solution highlights

Discuss Your AAV Characterization Study

Why AAV Characterization Is a Long-Range Sequence Problem

A recombinant AAV vector may be only a few kilobases long, yet the molecular population inside a vector preparation can be much more complex than a single expected reference sequence. Replication and packaging can generate incomplete genomes, asymmetric truncations, snapback or self-complementary structures, rearrangements, concatemers, reverse-packaged fragments, and DNA derived from production plasmids or host cells. A conventional coverage plot can show that all vector regions are represented while still obscuring how those regions are connected on individual molecules.

Long reads address that connectivity problem. When a read spans most or all of a packaged vector genome, the molecule can be interpreted as an intact structural unit rather than reconstructed from short fragments. This supports direct review of where molecules start and end, which vector elements remain linked, whether an unexpected fragment is fused to an ITR-bearing sequence, and whether distinct structural populations coexist in the same preparation. For upstream construct confirmation, our Full-Length Plasmid Sequencing service provides a complementary view of the DNA templates used during vector development.

The same long-range logic becomes even more important after transduction. Vector DNA in cells or tissues can exist as episomal forms, concatemers, truncated molecules, or host-integrated sequences. A host-vector junction may lie kilobases away from the transgene segment of interest, and an integration can contain multiple vector copies in different orientations. Targeted long-read or long-read whole-genome designs can retain these relationships on individual molecules when the relevant DNA is captured.

Long-read sequencing is therefore best viewed as a structural characterization layer. It complements, rather than replaces, assays for capsid concentration, vector genome titer, empty/full capsid ratio, potency, infectivity, residual DNA quantification, or validated release testing. The sequencing question is: what DNA structures are present, and how are their components connected?

What Can Long-Read Sequencing Characterize in an AAV Program?

Development questionLong-read evidenceResearch interpretation
Does packaged DNA match the intended vector genome?ITR-to-ITR or near-full-length read alignments, sequence identity, element order and orientationSupports review of expected vector-genome architecture and sequence discrepancies.
What heterogeneous genome forms are packaged?Read-level classification of complete, partial, truncated, rearranged, snapback/self-complementary, or concatemer-like structuresReveals structural populations that may be hidden by average coverage or bulk electrophoretic profiles.
Where do truncations or breakpoints recur?Read start/end positions, breakpoint clustering, element-level coverage transitionsIdentifies recurrent structural hotspots associated with a particular construct or preparation.
Are non-vector sequences present?Reads containing plasmid backbone, helper/packaging plasmid, or host-cell sequence when included in the reference search spaceProvides sequence-resolved evidence of captured DNA impurities or vector-linked chimeric molecules.
How is vector DNA organized after delivery?Long vector-containing reads, concatemer structures, episomal forms, host-vector junctionsSupports research into persistence, rearrangement, and integration architecture.
Where has AAV integrated?Host-vector junction reads, targeted enrichment or WGS evidence, breakpoint annotationMaps insertion loci and reconstructs local integration structures when coverage supports them.
Does methylation context matter?Native ONT DNA signal with methylation-aware calling in host-genome or vector-containing moleculesAdds an epigenetic layer when the biological question concerns chromatin or DNA modification.
How do lots, constructs, or time points differ?Matched structural classification and breakpoint summaries across samplesSupports comparative research without assuming fixed acceptance criteria.

The scope should be defined before sequencing. A purified vector lot requires a different workflow from a producer-cell genome, an animal tissue collected after dosing, or a plasmid template. We therefore design the sequencing strategy around the material and the decision rather than force every AAV project through one fixed pipeline.

AAV vector genome structural classes and long-read characterization scope

Packaged AAV Genome Characterization: Read the Molecule, Not Only the Coverage

Vector identity and full-length architecture

For purified rAAV preparations, long-read sequencing can be used to compare captured vector-derived molecules with the expected ITR-to-ITR design. Depending on sample preparation and platform, reads may span the complete vector genome or large fractions of it, allowing direct confirmation of element order, orientation, payload continuity, and sequence identity. PacBio HiFi is particularly useful when high consensus accuracy is important for sequence-level review, while ONT provides a flexible route for long-molecule profiling and can be advantageous when native-DNA information is part of the design.

Short-read sequencing remains valuable for high-depth base-level coverage, but fragmentation breaks the physical linkage among distant vector elements. For AAV, that linkage is often the point of the experiment. If a 5′ ITR-associated fragment, promoter, transgene segment, and plasmid-backbone sequence occur on one long molecule, that read provides structural information that four independent short-read coverage peaks cannot.

Partial genomes, truncations, and rearrangements

Long-read data can be classified at the read level to distinguish expected full-vector structures from partial or structurally altered forms. Useful outputs include length distributions, alignment start/end maps, breakpoint positions, orientation patterns, and molecule-level schematics. Recurrent start or end positions can reveal preferred truncation regions, while complex mappings can identify rearranged or concatemer-like molecules.

Published AAV studies demonstrate why this matters. Direct ITR-to-ITR nanopore sequencing has identified truncation hotspots in single-stranded and self-complementary vectors, while PacBio-based AAV genome population sequencing has shown that preparations appearing homogeneous by lower-resolution methods can contain substantial structural heterogeneity. These studies also show that platform and library-preparation biases must be considered when interpreting molecule frequencies. We therefore report read-supported structure with the relevant assay limitations rather than convert every read count into an absolute product-quality claim.

ITRs require explicit interpretation

Inverted terminal repeats are central to AAV replication and packaging and are also challenging sequencing features because of their secondary structure. The goal is not to claim that every ITR base is recovered perfectly in every workflow. Instead, we evaluate ITR-associated read starts and ends, ITR-spanning or ITR-adjacent sequence evidence, and structural continuity across the vector while documenting platform- and library-specific limitations. When an ITR sequence discrepancy could drive a development decision, we recommend orthogonal confirmation using a method designed for that specific question.

Residual and chimeric DNA sequence evidence

A long-read reference search can include the intended vector, production plasmids, helper or packaging constructs, and relevant host genomes. Reads that map partly or fully to these references can reveal plasmid-backbone fragments, rep/cap/helper-derived sequences, host-cell DNA fragments, or chimeric molecules that contain both AAV and non-AAV sequence. The analysis is sequence-specific: it tells you what captured molecules contain. It does not replace a validated quantitative residual-DNA assay when a program needs a formal concentration limit.

For production-template and cell-line questions that extend beyond packaged AAV DNA, our Viral Whole-Genome Resequencing capability and related long-read genomic workflows can be integrated into a broader characterization plan.

AAV Integration, Concatemers, and Vector Persistence

Characterizing packaged vector genomes and characterizing vector DNA after administration are related but distinct problems. A purified AAV preparation asks what is inside the capsid. A cell or tissue sample asks how vector-derived DNA is organized in a host-genome background, often at much lower abundance and with much more complex molecular structures.

Host-vector junction mapping

When high-molecular-weight genomic DNA is available, long reads can capture AAV sequence together with flanking host DNA. These chimeric reads provide direct junction evidence and can support mapping of integration loci, orientation of the inserted vector sequence, and the local host-genome context. If vector molecules are rare, targeted enrichment may be considered to increase the fraction of informative reads. For broader genome-scale studies, Human Whole Genome Sequencing can provide a complementary route to structural-variant and integration analysis.

Complex integrated structures

AAV integrations are not necessarily single, intact copies. Long molecules can contain head-to-tail, head-to-head, or tail-to-tail vector arrangements, mixtures of complete and truncated segments, and only one captured host junction when the complete integration exceeds the read length or enrichment window. Rather than forcing every event into a simple insertion model, our reporting distinguishes complete two-junction events from partially captured structures and preserves read-level evidence for review.

Targeted validation when the locus is known

When a candidate integration site or genomic locus has already been nominated, targeted long-range amplification or enrichment can concentrate coverage on that region and reconstruct junction architecture at higher depth. Our Human Long-Amplicon Sequencing capability can be used as a focused entry mode when the relevant locus is known and the main need is to resolve long-range allele structure.

Optional native methylation context

For selected host-genome or vector-containing DNA studies, ONT can retain native DNA modification information. This may be useful when integration-site chromatin, vector persistence, or host-locus methylation is part of the biological question. Our Long-Read Sequencing of DNA Methylation service provides additional detail on native methylation-aware workflows. Methylation analysis is optional and should not be added when it does not answer the development question.

PacBio HiFi, ONT, or a Combined AAV Strategy?

Decision factorPacBio HiFiOxford NanoporeCombined strategy
Primary strengthHigh consensus sequence accuracy with long-molecule structural contextFlexible long-read sequencing, native-DNA capability, rapid adaptation to targeted or genome-scale designsOrthogonal evidence when both high-accuracy structure and native/rapid long-read information are valuable
Packaged AAV genome profilingStrong for full/partial genome classification, breakpoint review, and sequence-level comparisonStrong for intact-vector profiling and structural heterogeneity, with platform-specific basecalling and length biases consideredUseful for cross-platform confirmation of major structural populations
ITR-focused interpretationHiFi consensus can support high-confidence sequence review; library handling and hairpin structure still matterCan sequence long native-derived molecules; ITR processivity/basecalling limitations must be evaluatedUseful when an ITR-related finding needs independent evidence
Host-vector integrationHigh-accuracy long reads for enriched or WGS-derived vector-containing moleculesLong native genomic molecules, targeted enrichment options, methylation-aware contextUseful for complex integration loci or development-stage confirmation
Native methylationNot the default readout for this solutionAvailable from native DNA when sample preparation preserves modification signalUse only when methylation is an explicit question
Best fitSequence-sensitive structural characterizationFlexible structural, native-DNA, and targeted investigationsHigh-value programs where complementary evidence justifies the added complexity

Our PacBio SMRT Sequencing Technology and Oxford Nanopore Sequencing Technology pages describe the broader platform principles. For this solution, platform choice is made from the AAV question, sample type, required sequence confidence, need for native-DNA information, and whether the study is targeted or genome-scale.

Integrated Gene Therapy and AAV Characterization Workflow

We combine the technical workflow and the project workflow so that every sequencing decision maps to a development question.

  1. Define the molecular question. Specify whether the project concerns production plasmid identity, packaged vector genomes, structural heterogeneity, residual/chimeric DNA, host integration, persistence, or methylation context.
  2. Review sample type and controls. Confirm vector design files, plasmid references, host-cell references, matched control material, and whether high-molecular-weight DNA or purified vector material is available.
  3. Select the assay scale. Choose packaged-vector profiling, targeted host-vector enrichment/amplicon analysis, or long-read WGS according to the expected abundance and genomic breadth of the event.
  4. Generate long-read data. Use PacBio HiFi, ONT, or a complementary strategy with project-specific library preparation and QC.
  5. Classify structures and junctions. Map reads to vector, plasmid, and host references; identify structural classes, truncation hotspots, chimeric molecules, integration sites, and read-level evidence.
  6. Interpret and prioritize findings. Deliver structure maps, quantitative summaries appropriate to the assay, limitations, and a prioritized list of events for orthogonal follow-up.

horizontal AAV gene therapy long-read characterization workflow

Bioinformatics Analysis and Evidence Reporting

Analysis moduleCore outputInterpretation notes
Read QC and reference mappingRead length/quality summaries, vector/plasmid/host alignment statisticsReference set is customized to the actual production system and study material.
Vector genome classificationExpected full-vector, partial/truncated, rearranged, snapback/self-complementary, concatemer-like, and other supported classesClassification rules are disclosed; ambiguous reads are retained rather than forced into a class.
Start/end and breakpoint profilingRead start/end density, recurrent breakpoint positions, structural hotspot plotsHighlights repeated packaging or truncation patterns without assuming mechanism from sequence alone.
Non-vector sequence screenReads mapping to plasmid backbone, helper/packaging constructs, or host genomeSequence evidence is not automatically equivalent to a validated impurity concentration.
Host-vector junction analysisIntegration coordinates, junction sequence, vector orientation, flanking host contextConfidence depends on unique mapping, read quality, breakpoint support, and assay breadth.
Concatemer/integration reconstructionMolecule schematics showing vector-copy order and orientationReports complete and partially captured events separately.
Comparative analysisLot/construct/time-point structural comparisons and prioritized differencesNo universal acceptance threshold is imposed unless supplied and justified by the sponsor.
Optional methylation analysisNative DNA methylation calls around vector-containing or host-genome regionsAvailable only when ONT-native DNA and suitable coverage support the question.

Every final report distinguishes direct read evidence, algorithmic classification, inferred interpretation, and items recommended for orthogonal validation. This is especially important for rare structural forms, ITR-associated discrepancies, and low-frequency integration events where library preparation, enrichment, and mapping can influence observed frequencies.

illustrative AAV long-read report with genome class breakpoint and host junction panels

How This Fits into Gene-Therapy Development

Regulatory guidance provides useful context for why complete vector sequence identity and characterization matter, but a research sequencing page should not imply that one assay constitutes a validated release method or satisfies a submission by itself. FDA's guidance on chemistry, manufacturing, and control information for human gene therapy INDs recommends full sequencing of viral vectors that are 40 kb or smaller and evaluation of discrepancies between expected and experimentally determined sequences. The same guidance also recognizes that AAV vectors are commonly produced by plasmid transfection and that sequencing material may be taken from drug substance or drug product when no master virus bank exists.

For an AAV development program, long reads can contribute sequence-resolved evidence in several places: confirming plasmid templates; characterizing packaged vector-genome structures; investigating unexpected bands or coverage patterns; identifying sequence-linked residual DNA; studying host-vector junctions in research models; and comparing structural profiles across development lots or vector designs. These data can complement orthogonal assays for identity, purity, capsid content, titer, potency, and residual process components.

What long-read sequencing does not establish by itself

  • Empty/full capsid ratio: a truly empty capsid contains no vector DNA to sequence. Capsid occupancy requires orthogonal physical or biochemical measurement.
  • Absolute residual-DNA specification: read counts can reveal sequence identity and relative representation but do not automatically replace a validated quantitative residual-DNA assay.
  • Potency or infectivity: genome integrity can inform mechanism and product understanding but does not measure functional transduction or expression by itself.
  • GMP release or regulatory validation: the service is Research Use Only; method validation, acceptance criteria, and regulatory use remain sponsor responsibilities.

Project Entry Modes and Sample Types

Entry materialTypical questionRecommended long-read mode
Production or transfer plasmid DNAIs the construct sequence and full plasmid architecture correct?Full-length plasmid sequencing with sequence and structural comparison.
Purified rAAV preparationWhat DNA structures are packaged, and are partial/rearranged or non-vector molecules present?Packaged vector-genome long-read profiling with customized reference search.
Producer-cell or cell-bank genomic DNAAre vector/plasmid sequences integrated, and what is their local architecture?Targeted or genome-scale long-read DNA sequencing depending on locus knowledge.
Transduced cells or preclinical tissue DNAHow does vector DNA persist, rearrange, concatemerize, or integrate?High-molecular-weight DNA with targeted enrichment or long-read WGS.
Known candidate integration locusWhat is the exact host-vector junction and copy orientation?Targeted long-amplicon or enrichment-based sequencing.
Native DNA for epigenetic researchIs methylation around vector-containing or host loci relevant?ONT native-DNA sequencing with methylation-aware analysis.

Exact input quantity and quality requirements depend on platform, assay scale, and whether native high-molecular-weight DNA must be preserved. We review sample requirements before project initiation rather than publish one universal threshold for all AAV workflows. General shipment and preparation guidance is available in our Sample Submission Guideline.

Why Choose CD Genomics for AAV Long-Read Characterization?

PacBio and ONT under one project strategy

We do not force AAV studies onto a single long-read platform. PacBio HiFi and ONT provide different combinations of consensus accuracy, native-DNA capability, targeted flexibility, and genome-scale context. We choose the workflow from the structural question and can use complementary evidence when one platform alone would leave an important uncertainty.

From production construct to host-genome context

The same program may need full plasmid confirmation, packaged AAV genome profiling, host-vector integration mapping, and later genome-scale or methylation-aware investigation. Our long-read portfolio allows these questions to be handled as a connected characterization program rather than unrelated sequencing orders.

Read-level evidence instead of only summary percentages

We provide molecule-level structure diagrams, breakpoint evidence, mapping context, and classification rules so that a reported structural population can be traced back to representative reads. This is important for complex vector forms and integration events that cannot be understood from one aggregate number.

Explicit limitations and orthogonal follow-up

AAV measurements are vulnerable to biases from DNA extraction, library preparation, secondary structure, enrichment, read length, and reference mapping. We document these limitations and identify which findings should be confirmed by an independent method rather than present every sequence observation as a definitive quality attribute.

Independent Published Example: Long-Read Reconstruction of AAV Genomes and Integration Architecture

Greig JA, Martins KM, Breton C, et al. Integrated vector genomes may contribute to long-term expression in primate liver after AAV administration. Nature Biotechnology. 2024;42:1232–1242. doi:10.1038/s41587-023-01974-7.

Background

Greig and colleagues investigated the persistence and expression of liver-directed AAV8 and AAVrh10 vectors in nonhuman primates over extended follow-up. The study combined molecular, histologic, transcriptomic, integration-site, and long-read approaches to understand why vector DNA could persist even as transgene expression declined. This makes the work relevant to gene-therapy characterization because it separates the structure of administered vector preparations from the structure of vector DNA that remained in tissues.

Methods

For late tissue samples, the authors enriched vector-containing high-molecular-weight DNA with probes tiled across vector transcriptional units and performed PacBio Sequel II HiFi circular consensus sequencing. Reported CCS reads ranged from 51 to 50,419 bp, with average sequence sizes across samples ranging from 4,216 to 6,113 bp. Reads containing vector sequence were mapped against both vector and host genomes to characterize vector organization and host-genome junctions.

Results

The study first showed that complete ITR-to-ITR vector genomes represented the majority of DNA in the vector preparations used for dosing, while a minority population contained truncations that were not present in the input plasmid. In late liver samples, long-read analysis revealed extensive heterogeneity, including rearranged and truncated vector segments organized into complex concatemers. The authors observed head-to-tail, head-to-head, and tail-to-tail concatemeric configurations and captured host-vector junctions within long individual molecules.

Integration-site analysis showed that vector integration events persisted at broadly distributed loci. At later time points, reported integration levels stabilized in the range of 0.1–0.7 AAV integration events per 100 genomes in the evaluated samples. The authors also reported that the identified integration sites in these nonhuman primates were not located in the genic regions they evaluated as frequently mutated in human hepatocellular carcinoma or previously implicated in AAV-associated mouse liver tumors.

published Figure 5 showing complex integrated AAV vector DNA structures in nonhuman primate liverFigure 5 from Greig et al. (Nature Biotechnology, 2024, CC BY 4.0) shows representative structures of integrated vector DNA after AAV administration, including complex concatemeric configurations and host genomic junctions.

Conclusion

This independent published example illustrates two distinct uses of long-read sequencing in an AAV program: profiling the integrity of the administered vector genome and reconstructing the complex organization of vector DNA in host tissue. It also shows why the expected plasmid or packaged-vector sequence is not sufficient to predict every structure that may arise after delivery. The findings are presented here as literature evidence, not as a CD Genomics customer project or a performance guarantee.

FAQs

Sample Deliverables

1. AAV genome integrity overview — read-level classification of expected full-vector structures, partial/truncated molecules, rearrangements, concatemer-like forms, and other supported categories.

2. Read start/end and breakpoint map — recurrent vector-genome start positions, end positions, structural breakpoints, and hotspots mapped against annotated vector elements.

3. Non-vector sequence report — reads mapping to production plasmids, helper/packaging sequences, host genomes, and vector-linked chimeric molecules.

4. Integration-site and architecture report — host-vector junction coordinates, surrounding host sequence, vector-copy orientation, concatemer structure, and representative long reads.

5. Comparative construct/lot summary — consistent structural metrics and plots for side-by-side review across vector designs, lots, or research time points.

sample AAV long-read characterization deliverables with genome classes breakpoints and integration map

References

  1. Greig JA, Martins KM, Breton C, et al. Integrated vector genomes may contribute to long-term expression in primate liver after AAV administration. Nature Biotechnology. 2024;42:1232–1242. doi:10.1038/s41587-023-01974-7.
  2. Namkung S, Tran NT, Manokaran S, et al. Direct ITR-to-ITR Nanopore Sequencing of AAV Vector Genomes. Human Gene Therapy. 2022;33:1187–1196. doi:10.1089/hum.2022.143.
  3. Kosaka M, Fujita Sajiki A, Fujita K, et al. Evaluation of the loading capacity and patterns of packaged DNA in AAV genomes of different sizes using long-read sequencing. Molecular Therapy: Methods & Clinical Development. 2025;33:101474. doi:10.1016/j.omtm.2025.101474.
  4. U.S. Food and Drug Administration. Chemistry, Manufacturing, and Control Information for Human Gene Therapy Investigational New Drug Applications: Guidance for Industry.

For Research Use Only. Not for use in diagnostic or clinical procedures.

Get Your Instant Quote