
Gene-therapy vector characterization becomes difficult when the question is not simply whether a transgene sequence is present, but whether the complete vector genome is packaged as designed, which partial or rearranged forms coexist in the preparation, whether plasmid or host-cell sequences are present, and how vector DNA is organized after delivery. These are fundamentally long-range sequence questions.
CD Genomics provides long-read sequencing for gene therapy and AAV characterization using complementary PacBio HiFi and Oxford Nanopore Technologies (ONT) workflows. We help research and biopharma teams characterize packaged recombinant AAV (rAAV) genomes, plasmid templates, complex vector forms, host-vector junctions, integration architecture, and selected epigenetic features with read-level structural evidence that can span entire vector molecules or long host-genome contexts.
Solution highlights
A recombinant AAV vector may be only a few kilobases long, yet the molecular population inside a vector preparation can be much more complex than a single expected reference sequence. Replication and packaging can generate incomplete genomes, asymmetric truncations, snapback or self-complementary structures, rearrangements, concatemers, reverse-packaged fragments, and DNA derived from production plasmids or host cells. A conventional coverage plot can show that all vector regions are represented while still obscuring how those regions are connected on individual molecules.
Long reads address that connectivity problem. When a read spans most or all of a packaged vector genome, the molecule can be interpreted as an intact structural unit rather than reconstructed from short fragments. This supports direct review of where molecules start and end, which vector elements remain linked, whether an unexpected fragment is fused to an ITR-bearing sequence, and whether distinct structural populations coexist in the same preparation. For upstream construct confirmation, our Full-Length Plasmid Sequencing service provides a complementary view of the DNA templates used during vector development.
The same long-range logic becomes even more important after transduction. Vector DNA in cells or tissues can exist as episomal forms, concatemers, truncated molecules, or host-integrated sequences. A host-vector junction may lie kilobases away from the transgene segment of interest, and an integration can contain multiple vector copies in different orientations. Targeted long-read or long-read whole-genome designs can retain these relationships on individual molecules when the relevant DNA is captured.
Long-read sequencing is therefore best viewed as a structural characterization layer. It complements, rather than replaces, assays for capsid concentration, vector genome titer, empty/full capsid ratio, potency, infectivity, residual DNA quantification, or validated release testing. The sequencing question is: what DNA structures are present, and how are their components connected?
| Development question | Long-read evidence | Research interpretation |
| Does packaged DNA match the intended vector genome? | ITR-to-ITR or near-full-length read alignments, sequence identity, element order and orientation | Supports review of expected vector-genome architecture and sequence discrepancies. |
| What heterogeneous genome forms are packaged? | Read-level classification of complete, partial, truncated, rearranged, snapback/self-complementary, or concatemer-like structures | Reveals structural populations that may be hidden by average coverage or bulk electrophoretic profiles. |
| Where do truncations or breakpoints recur? | Read start/end positions, breakpoint clustering, element-level coverage transitions | Identifies recurrent structural hotspots associated with a particular construct or preparation. |
| Are non-vector sequences present? | Reads containing plasmid backbone, helper/packaging plasmid, or host-cell sequence when included in the reference search space | Provides sequence-resolved evidence of captured DNA impurities or vector-linked chimeric molecules. |
| How is vector DNA organized after delivery? | Long vector-containing reads, concatemer structures, episomal forms, host-vector junctions | Supports research into persistence, rearrangement, and integration architecture. |
| Where has AAV integrated? | Host-vector junction reads, targeted enrichment or WGS evidence, breakpoint annotation | Maps insertion loci and reconstructs local integration structures when coverage supports them. |
| Does methylation context matter? | Native ONT DNA signal with methylation-aware calling in host-genome or vector-containing molecules | Adds an epigenetic layer when the biological question concerns chromatin or DNA modification. |
| How do lots, constructs, or time points differ? | Matched structural classification and breakpoint summaries across samples | Supports comparative research without assuming fixed acceptance criteria. |
The scope should be defined before sequencing. A purified vector lot requires a different workflow from a producer-cell genome, an animal tissue collected after dosing, or a plasmid template. We therefore design the sequencing strategy around the material and the decision rather than force every AAV project through one fixed pipeline.

For purified rAAV preparations, long-read sequencing can be used to compare captured vector-derived molecules with the expected ITR-to-ITR design. Depending on sample preparation and platform, reads may span the complete vector genome or large fractions of it, allowing direct confirmation of element order, orientation, payload continuity, and sequence identity. PacBio HiFi is particularly useful when high consensus accuracy is important for sequence-level review, while ONT provides a flexible route for long-molecule profiling and can be advantageous when native-DNA information is part of the design.
Short-read sequencing remains valuable for high-depth base-level coverage, but fragmentation breaks the physical linkage among distant vector elements. For AAV, that linkage is often the point of the experiment. If a 5′ ITR-associated fragment, promoter, transgene segment, and plasmid-backbone sequence occur on one long molecule, that read provides structural information that four independent short-read coverage peaks cannot.
Long-read data can be classified at the read level to distinguish expected full-vector structures from partial or structurally altered forms. Useful outputs include length distributions, alignment start/end maps, breakpoint positions, orientation patterns, and molecule-level schematics. Recurrent start or end positions can reveal preferred truncation regions, while complex mappings can identify rearranged or concatemer-like molecules.
Published AAV studies demonstrate why this matters. Direct ITR-to-ITR nanopore sequencing has identified truncation hotspots in single-stranded and self-complementary vectors, while PacBio-based AAV genome population sequencing has shown that preparations appearing homogeneous by lower-resolution methods can contain substantial structural heterogeneity. These studies also show that platform and library-preparation biases must be considered when interpreting molecule frequencies. We therefore report read-supported structure with the relevant assay limitations rather than convert every read count into an absolute product-quality claim.
Inverted terminal repeats are central to AAV replication and packaging and are also challenging sequencing features because of their secondary structure. The goal is not to claim that every ITR base is recovered perfectly in every workflow. Instead, we evaluate ITR-associated read starts and ends, ITR-spanning or ITR-adjacent sequence evidence, and structural continuity across the vector while documenting platform- and library-specific limitations. When an ITR sequence discrepancy could drive a development decision, we recommend orthogonal confirmation using a method designed for that specific question.
A long-read reference search can include the intended vector, production plasmids, helper or packaging constructs, and relevant host genomes. Reads that map partly or fully to these references can reveal plasmid-backbone fragments, rep/cap/helper-derived sequences, host-cell DNA fragments, or chimeric molecules that contain both AAV and non-AAV sequence. The analysis is sequence-specific: it tells you what captured molecules contain. It does not replace a validated quantitative residual-DNA assay when a program needs a formal concentration limit.
For production-template and cell-line questions that extend beyond packaged AAV DNA, our Viral Whole-Genome Resequencing capability and related long-read genomic workflows can be integrated into a broader characterization plan.
Characterizing packaged vector genomes and characterizing vector DNA after administration are related but distinct problems. A purified AAV preparation asks what is inside the capsid. A cell or tissue sample asks how vector-derived DNA is organized in a host-genome background, often at much lower abundance and with much more complex molecular structures.
When high-molecular-weight genomic DNA is available, long reads can capture AAV sequence together with flanking host DNA. These chimeric reads provide direct junction evidence and can support mapping of integration loci, orientation of the inserted vector sequence, and the local host-genome context. If vector molecules are rare, targeted enrichment may be considered to increase the fraction of informative reads. For broader genome-scale studies, Human Whole Genome Sequencing can provide a complementary route to structural-variant and integration analysis.
AAV integrations are not necessarily single, intact copies. Long molecules can contain head-to-tail, head-to-head, or tail-to-tail vector arrangements, mixtures of complete and truncated segments, and only one captured host junction when the complete integration exceeds the read length or enrichment window. Rather than forcing every event into a simple insertion model, our reporting distinguishes complete two-junction events from partially captured structures and preserves read-level evidence for review.
When a candidate integration site or genomic locus has already been nominated, targeted long-range amplification or enrichment can concentrate coverage on that region and reconstruct junction architecture at higher depth. Our Human Long-Amplicon Sequencing capability can be used as a focused entry mode when the relevant locus is known and the main need is to resolve long-range allele structure.
For selected host-genome or vector-containing DNA studies, ONT can retain native DNA modification information. This may be useful when integration-site chromatin, vector persistence, or host-locus methylation is part of the biological question. Our Long-Read Sequencing of DNA Methylation service provides additional detail on native methylation-aware workflows. Methylation analysis is optional and should not be added when it does not answer the development question.
| Decision factor | PacBio HiFi | Oxford Nanopore | Combined strategy |
| Primary strength | High consensus sequence accuracy with long-molecule structural context | Flexible long-read sequencing, native-DNA capability, rapid adaptation to targeted or genome-scale designs | Orthogonal evidence when both high-accuracy structure and native/rapid long-read information are valuable |
| Packaged AAV genome profiling | Strong for full/partial genome classification, breakpoint review, and sequence-level comparison | Strong for intact-vector profiling and structural heterogeneity, with platform-specific basecalling and length biases considered | Useful for cross-platform confirmation of major structural populations |
| ITR-focused interpretation | HiFi consensus can support high-confidence sequence review; library handling and hairpin structure still matter | Can sequence long native-derived molecules; ITR processivity/basecalling limitations must be evaluated | Useful when an ITR-related finding needs independent evidence |
| Host-vector integration | High-accuracy long reads for enriched or WGS-derived vector-containing molecules | Long native genomic molecules, targeted enrichment options, methylation-aware context | Useful for complex integration loci or development-stage confirmation |
| Native methylation | Not the default readout for this solution | Available from native DNA when sample preparation preserves modification signal | Use only when methylation is an explicit question |
| Best fit | Sequence-sensitive structural characterization | Flexible structural, native-DNA, and targeted investigations | High-value programs where complementary evidence justifies the added complexity |
Our PacBio SMRT Sequencing Technology and Oxford Nanopore Sequencing Technology pages describe the broader platform principles. For this solution, platform choice is made from the AAV question, sample type, required sequence confidence, need for native-DNA information, and whether the study is targeted or genome-scale.
We combine the technical workflow and the project workflow so that every sequencing decision maps to a development question.

| Analysis module | Core output | Interpretation notes |
| Read QC and reference mapping | Read length/quality summaries, vector/plasmid/host alignment statistics | Reference set is customized to the actual production system and study material. |
| Vector genome classification | Expected full-vector, partial/truncated, rearranged, snapback/self-complementary, concatemer-like, and other supported classes | Classification rules are disclosed; ambiguous reads are retained rather than forced into a class. |
| Start/end and breakpoint profiling | Read start/end density, recurrent breakpoint positions, structural hotspot plots | Highlights repeated packaging or truncation patterns without assuming mechanism from sequence alone. |
| Non-vector sequence screen | Reads mapping to plasmid backbone, helper/packaging constructs, or host genome | Sequence evidence is not automatically equivalent to a validated impurity concentration. |
| Host-vector junction analysis | Integration coordinates, junction sequence, vector orientation, flanking host context | Confidence depends on unique mapping, read quality, breakpoint support, and assay breadth. |
| Concatemer/integration reconstruction | Molecule schematics showing vector-copy order and orientation | Reports complete and partially captured events separately. |
| Comparative analysis | Lot/construct/time-point structural comparisons and prioritized differences | No universal acceptance threshold is imposed unless supplied and justified by the sponsor. |
| Optional methylation analysis | Native DNA methylation calls around vector-containing or host-genome regions | Available only when ONT-native DNA and suitable coverage support the question. |
Every final report distinguishes direct read evidence, algorithmic classification, inferred interpretation, and items recommended for orthogonal validation. This is especially important for rare structural forms, ITR-associated discrepancies, and low-frequency integration events where library preparation, enrichment, and mapping can influence observed frequencies.

Regulatory guidance provides useful context for why complete vector sequence identity and characterization matter, but a research sequencing page should not imply that one assay constitutes a validated release method or satisfies a submission by itself. FDA's guidance on chemistry, manufacturing, and control information for human gene therapy INDs recommends full sequencing of viral vectors that are 40 kb or smaller and evaluation of discrepancies between expected and experimentally determined sequences. The same guidance also recognizes that AAV vectors are commonly produced by plasmid transfection and that sequencing material may be taken from drug substance or drug product when no master virus bank exists.
For an AAV development program, long reads can contribute sequence-resolved evidence in several places: confirming plasmid templates; characterizing packaged vector-genome structures; investigating unexpected bands or coverage patterns; identifying sequence-linked residual DNA; studying host-vector junctions in research models; and comparing structural profiles across development lots or vector designs. These data can complement orthogonal assays for identity, purity, capsid content, titer, potency, and residual process components.
| Entry material | Typical question | Recommended long-read mode |
| Production or transfer plasmid DNA | Is the construct sequence and full plasmid architecture correct? | Full-length plasmid sequencing with sequence and structural comparison. |
| Purified rAAV preparation | What DNA structures are packaged, and are partial/rearranged or non-vector molecules present? | Packaged vector-genome long-read profiling with customized reference search. |
| Producer-cell or cell-bank genomic DNA | Are vector/plasmid sequences integrated, and what is their local architecture? | Targeted or genome-scale long-read DNA sequencing depending on locus knowledge. |
| Transduced cells or preclinical tissue DNA | How does vector DNA persist, rearrange, concatemerize, or integrate? | High-molecular-weight DNA with targeted enrichment or long-read WGS. |
| Known candidate integration locus | What is the exact host-vector junction and copy orientation? | Targeted long-amplicon or enrichment-based sequencing. |
| Native DNA for epigenetic research | Is methylation around vector-containing or host loci relevant? | ONT native-DNA sequencing with methylation-aware analysis. |
Exact input quantity and quality requirements depend on platform, assay scale, and whether native high-molecular-weight DNA must be preserved. We review sample requirements before project initiation rather than publish one universal threshold for all AAV workflows. General shipment and preparation guidance is available in our Sample Submission Guideline.
We do not force AAV studies onto a single long-read platform. PacBio HiFi and ONT provide different combinations of consensus accuracy, native-DNA capability, targeted flexibility, and genome-scale context. We choose the workflow from the structural question and can use complementary evidence when one platform alone would leave an important uncertainty.
The same program may need full plasmid confirmation, packaged AAV genome profiling, host-vector integration mapping, and later genome-scale or methylation-aware investigation. Our long-read portfolio allows these questions to be handled as a connected characterization program rather than unrelated sequencing orders.
We provide molecule-level structure diagrams, breakpoint evidence, mapping context, and classification rules so that a reported structural population can be traced back to representative reads. This is important for complex vector forms and integration events that cannot be understood from one aggregate number.
AAV measurements are vulnerable to biases from DNA extraction, library preparation, secondary structure, enrichment, read length, and reference mapping. We document these limitations and identify which findings should be confirmed by an independent method rather than present every sequence observation as a definitive quality attribute.
Greig JA, Martins KM, Breton C, et al. Integrated vector genomes may contribute to long-term expression in primate liver after AAV administration. Nature Biotechnology. 2024;42:1232–1242. doi:10.1038/s41587-023-01974-7.
Greig and colleagues investigated the persistence and expression of liver-directed AAV8 and AAVrh10 vectors in nonhuman primates over extended follow-up. The study combined molecular, histologic, transcriptomic, integration-site, and long-read approaches to understand why vector DNA could persist even as transgene expression declined. This makes the work relevant to gene-therapy characterization because it separates the structure of administered vector preparations from the structure of vector DNA that remained in tissues.
For late tissue samples, the authors enriched vector-containing high-molecular-weight DNA with probes tiled across vector transcriptional units and performed PacBio Sequel II HiFi circular consensus sequencing. Reported CCS reads ranged from 51 to 50,419 bp, with average sequence sizes across samples ranging from 4,216 to 6,113 bp. Reads containing vector sequence were mapped against both vector and host genomes to characterize vector organization and host-genome junctions.
The study first showed that complete ITR-to-ITR vector genomes represented the majority of DNA in the vector preparations used for dosing, while a minority population contained truncations that were not present in the input plasmid. In late liver samples, long-read analysis revealed extensive heterogeneity, including rearranged and truncated vector segments organized into complex concatemers. The authors observed head-to-tail, head-to-head, and tail-to-tail concatemeric configurations and captured host-vector junctions within long individual molecules.
Integration-site analysis showed that vector integration events persisted at broadly distributed loci. At later time points, reported integration levels stabilized in the range of 0.1–0.7 AAV integration events per 100 genomes in the evaluated samples. The authors also reported that the identified integration sites in these nonhuman primates were not located in the genic regions they evaluated as frequently mutated in human hepatocellular carcinoma or previously implicated in AAV-associated mouse liver tumors.
Figure 5 from Greig et al. (Nature Biotechnology, 2024, CC BY 4.0) shows representative structures of integrated vector DNA after AAV administration, including complex concatemeric configurations and host genomic junctions.
This independent published example illustrates two distinct uses of long-read sequencing in an AAV program: profiling the integrity of the administered vector genome and reconstructing the complex organization of vector DNA in host tissue. It also shows why the expected plasmid or packaged-vector sequence is not sufficient to predict every structure that may arise after delivery. The findings are presented here as literature evidence, not as a CD Genomics customer project or a performance guarantee.
The main advantage is molecule-level structural context. A long read can connect distant vector elements, ITR-associated ends, rearrangements, plasmid or host sequence, and host-vector junctions on the same molecule. This helps distinguish complete vector genomes from partial or structurally complex forms that may look similar in average coverage.
No. A truly empty capsid contains no vector DNA to sequence. Long-read sequencing characterizes DNA-containing molecules recovered from the preparation. Empty/full capsid ratio requires an orthogonal capsid-content method. We therefore avoid interpreting sequencing-derived full/partial genome percentages as the physical fraction of all capsids in a sample.
Yes, when those structures are represented in captured sequencing molecules. We classify reads according to their alignment pattern, start and end positions, breakpoints, orientation, and vector-element composition. Frequencies are interpreted within the context of the library and sequencing workflow because extraction, library preparation, and read-length bias can influence observed representation.
Long-read reference mapping can identify captured molecules derived from plasmid backbone, packaging/helper constructs, or the host genome and can reveal chimeric molecules when non-vector sequence is linked to AAV sequence. This is sequence-resolved evidence, but it does not automatically replace a validated quantitative residual-DNA assay required for a formal specification.
Yes, when vector-containing DNA is captured with sufficient long-range context. We can map host-vector junctions and reconstruct local integration architecture using targeted enrichment, targeted long-amplicon designs, or long-read WGS. The optimal strategy depends on vector abundance, whether candidate loci are already known, DNA quality, and the breadth of the question.
PacBio HiFi is attractive when high consensus sequence accuracy and read-level structural classification are the priority. ONT is attractive for flexible long-molecule and native-DNA workflows, including optional methylation context. For high-value studies, complementary use can help confirm major structural observations. We select the platform after reviewing the vector design and development question.
No. This is a Research Use Only sequencing and characterization service. The resulting data may support product understanding and development studies, but method validation, specifications, acceptance criteria, GMP release use, and regulatory submission strategy remain the sponsor's responsibility.
1. AAV genome integrity overview — read-level classification of expected full-vector structures, partial/truncated molecules, rearrangements, concatemer-like forms, and other supported categories.
2. Read start/end and breakpoint map — recurrent vector-genome start positions, end positions, structural breakpoints, and hotspots mapped against annotated vector elements.
3. Non-vector sequence report — reads mapping to production plasmids, helper/packaging sequences, host genomes, and vector-linked chimeric molecules.
4. Integration-site and architecture report — host-vector junction coordinates, surrounding host sequence, vector-copy orientation, concatemer structure, and representative long reads.
5. Comparative construct/lot summary — consistent structural metrics and plots for side-by-side review across vector designs, lots, or research time points.

References
For Research Use Only. Not for use in diagnostic or clinical procedures.