Long-Read Sequencing for Gene-Editing Characterization

Long-Read Sequencing for Gene-Editing Characterization

long-read sequencing for gene-editing characterization and genome integrity research

A successful gene-editing experiment is not defined only by the percentage of reads carrying the intended small edit. Nuclease cutting, donor-template repair, base editing, and prime editing can create a mixture of alleles that includes the desired sequence together with larger deletions, insertions, inversions, rearrangements, donor or vector integrations, and allele-specific outcomes that short amplicons may not fully resolve.

CD Genomics provides long-read sequencing for gene-editing characterization across targeted and genome-scale research designs. Using complementary PacBio HiFi and Oxford Nanopore Technologies (ONT) workflows, we help research and biopharma teams define on-target architecture, resolve long-range editing alleles, phase intended and unintended sequence changes, characterize knock-in structure, investigate nominated off-target loci, and assess broader genome-integrity questions when whole-genome analysis is warranted.

Solution highlights

Discuss Your Gene-Editing Study

Why Editing Efficiency Alone Is Not Full Characterization

Short amplicon sequencing is highly effective for measuring small insertions and deletions close to a programmed cut site. Sanger sequencing remains useful for confirming a defined clonal sequence. These approaches answer an important question: did editing occur at the expected local sequence? They do not always answer a second, broader question: what complete molecular structure was produced by the editing and repair process?

Double-strand-break repair can extend beyond the few hundred bases typically captured in a short amplicon. Published studies have documented kilobase-scale deletions, long insertions, inversions, translocations, complex local rearrangements, and incorporation of exogenous DNA after CRISPR editing. Even editing systems designed to avoid a conventional double-strand break may generate a spectrum of intended and unintended local outcomes that must be interpreted in the context of the exact editor, delivery system, cell type, target locus, and repair pathway.

Long-read sequencing adds value when the structure of the edited allele is itself part of the answer. A single molecule can span the intended edit, adjacent variants, donor junctions, repeated sequences, or a long deletion breakpoint. This reduces dependence on assembling an editing outcome from multiple disconnected short fragments and makes it possible to separate distinct alleles in mixed or mosaic samples. Our Human Long-Amplicon Sequencing capability provides a targeted entry point when the locus is known and the principal question is the complete spectrum of local alleles.

Long reads are not a universal replacement for short-read NGS, ddPCR, cytogenetics, optical mapping, or dedicated genome-wide off-target nomination assays. Instead, we use them as a high-information characterization layer. The study design is built around the failure mode that matters: local structural complexity, allelic dropout, donor integration, long-range phasing, genome-wide structural change, or the need to connect several observations on the same DNA molecule.

What Gene-Editing Outcomes Can Be Characterized?

Research questionLong-read evidenceWhat it helps determine
Was the intended edit introduced correctly?Full target-spanning reads and allele-specific sequence comparisonConfirms intended sequence change while retaining long-range context around the edit.
Are larger deletions or insertions present?Breakpoint-spanning reads, long alignments, local assemblyDetects events that extend beyond a conventional short amplicon and reconstructs their boundaries.
Did the donor or delivery vector integrate unexpectedly?Genome-donor or genome-vector junctions and inserted-sequence structureDistinguishes intended knock-in architecture from partial, tandem, inverted, backbone, or other exogenous integrations.
Are multiple editing outcomes present in the sample?Read-level allele clustering and haplotype reconstructionSeparates wild-type, intended, partially edited, mosaic, and complex structural alleles when coverage supports them.
Which chromosome carries the intended edit?Phasing of the edit with nearby heterozygous variantsSupports allele-specific interpretation and helps identify compound or asymmetric editing outcomes.
What happened at nominated off-target sites?Long target-spanning reads for pre-specified candidate lociCharacterizes small and structural editing outcomes at loci nominated by the sponsor or another assay.
Is there evidence of broader genome-integrity change?Long-read WGS, structural-variant calling, copy-number and breakpoint reviewExpands beyond selected loci to evaluate larger rearrangements and unexpected structural events.
Does epigenetic context matter?Native ONT signal with methylation-aware analysis when appropriateAdds methylation state to the same long molecules used for sequence and structural analysis.

The appropriate scope depends on the editing modality. A simple knockout clone may need only a long target amplicon. A knock-in program using a donor plasmid or viral vector may require explicit analysis of donor-host junctions and unexpected exogenous integrations. A therapeutic research program with a genome-integrity question may justify a broader long-read whole-genome strategy.

gene editing outcome spectrum from intended edit to complex structural variants

On-Target Characterization: Reconstruct the Edited Allele, Not Just the Cut Site

Small edits in long-range haplotype context

For nuclease knockouts, base editing, prime editing, or precise sequence correction, the intended nucleotide-level event may be small while the surrounding allele is not. Long reads can phase the intended edit with nearby heterozygous variants and distinguish separate alleles in a mixed population. This is particularly useful when apparent homozygosity could instead reflect allelic dropout, a large deletion on one chromosome, or unequal representation of different edited alleles.

Large deletions, inversions, and complex rearrangements

Large structural outcomes can escape assays designed around a narrow expected product. We examine long alignments, split reads, breakpoint-spanning molecules, and local assemblies to identify deletions, insertions, inversions, duplications, and complex combinations supported by the sequencing data. For projects centered on these events, our Human Genome Structural Variation Detection workflow can extend the analysis beyond the immediate edit site.

Knock-in and donor-template architecture

For HDR-mediated or other targeted integration designs, a positive junction PCR does not necessarily establish the complete integrated structure. Long reads can determine whether the donor is present in the intended orientation, whether both junctions are correct, whether partial donor fragments or vector backbone sequences are present, and whether multiple donor copies are arranged in tandem or more complex configurations. These questions are especially important when an apparently successful edit must be reduced to a discrete, reviewable allele model.

Clonal versus pooled samples

Clonal samples simplify interpretation but do not eliminate the possibility of unexpected structural alleles. Pooled edited populations can contain many outcomes simultaneously. We therefore distinguish between allele discovery and allele frequency estimation. Read counts can support relative abundance estimates, but PCR amplification, DNA quality, fragment length, and enrichment method can bias recovery of long or structurally unusual molecules. Where precise quantitation is critical, we recommend integrating long-read structural evidence with an orthogonal quantitative assay.

Off-Target Confirmation and Genome-Integrity Assessment

Genome-editing safety analysis is often divided into two distinct tasks: off-target site nomination and off-target site confirmation/characterization. Long-read sequencing can be highly informative for the second task, because it can show the size and structure of an editing event once a candidate locus is known. It should not be presented as a universal substitute for every biochemical, cell-based, in silico, or genome-wide nomination method.

For sponsor-nominated candidate sites, targeted long-read sequencing can characterize both small sequence changes and larger structural outcomes. This matters because an off-target cut that produces a several-hundred-base deletion or rearrangement may not be adequately described by a short assay centered tightly on the predicted cleavage coordinate.

When the research question is broader than a defined candidate list, Human Whole Genome Sequencing with long reads can be used to examine structural variation and long-range genome architecture across the genome. Whole-genome data may be combined with long-read variant calling to review SNVs/indels, larger SVs, breakpoint structure, and haplotype context. The analytical plan should include an unedited or parental control whenever possible so that pre-existing variants are not incorrectly attributed to editing.

For vector-delivered editing systems, genome-integrity analysis may also include searches for exogenous sequence integrations. The interpretation should distinguish intended donor sequence from vector backbone, scaffold, viral-vector fragments, or other inserted sequence and should report the evidence supporting each junction rather than reducing all insertions to a single count.

PacBio HiFi, ONT, or a Combined Strategy?

StrategyBest fitKey strengthImportant consideration
PacBio HiFi targeted long readsHigh-confidence local allele reconstruction, knock-in verification, phased sequence analysisHigh per-read consensus accuracy supports detailed sequence-level comparison across long amplicons or enriched regions.Targeted designs remain bounded by the region recovered and can be affected by amplification or enrichment bias.
ONT targeted long readsVery long target regions, complex insertions, repetitive loci, structural architectureLong molecules can span large events and connect multiple junctions; native DNA can retain methylation signal.Read-level sequence errors and coverage variability require appropriate consensus and variant-support thresholds.
PacBio HiFi long-read WGSBroad genome-integrity assessment with accurate SV and small-variant analysisCombines long-range structural evidence with high sequence accuracy and phasing.Whole-genome studies require more DNA, sequencing, and analysis than a focused target assay.
ONT long-read WGSUltra-long structural context, complex rearrangements, native methylation-aware analysisVery long molecules can span large repeats, rearrangements, and inserted sequence while retaining native signal.Study design should explicitly define coverage and variant evidence appropriate to the biological question.
Combined long-read + orthogonal assaysPrograms requiring both structural reconstruction and independent quantitative confirmationSeparates structure, abundance, and functional questions rather than forcing one method to answer all three.Cross-platform discrepancies should be resolved using pre-defined interpretation rules.

We select the platform from the characterization objective rather than from a fixed technology preference. PacBio SMRT sequencing is well suited to high-accuracy allele reconstruction, while Oxford Nanopore sequencing offers very long native DNA reads and optional methylation-aware analysis. Some projects benefit from using one platform for discovery and an orthogonal method for confirmation.

Integrated Gene-Editing Characterization Workflow

Our workflow begins with the edit design and the risk question, not with a preselected sequencing kit. The same CRISPR edit can require very different assays depending on whether the goal is local clone verification, donor integration characterization, nominated off-target confirmation, or broader genome-integrity evaluation.

horizontal workflow for long-read gene-editing characterization

1. Define the expected edit and plausible failure modes

We review the target locus, guide or editing coordinates, editor type, donor/vector sequence if used, delivery method, expected allele, known polymorphisms, and any prior short-read, Sanger, ddPCR, or cytogenetic results. The goal is to define what must be observed directly and what could cause a false sense of a successful edit.

2. Match the assay scale to the question

A long amplicon may be sufficient for a focused on-target study. A targeted amplification-free design can be considered when PCR bias or allelic dropout is a concern. Whole-genome long reads are considered when large rearrangements, unknown insertion sites, or broader genome integrity are the research question.

3. Sequence edited and relevant control material

Whenever possible, edited material is analyzed alongside an unedited parental, donor-matched, or process-relevant control. Controls are critical for separating editing-associated changes from pre-existing structural variants, cell-line background, or sequencing artifacts.

4. Reconstruct alleles and structural events

Reads are aligned to the reference genome and, when relevant, to donor or vector sequences. We cluster alleles, identify breakpoint-spanning reads, reconstruct local structures, phase variants, and evaluate support for each candidate event.

5. Compare across loci, clones, conditions, or passages

For clone selection or process-development studies, we compare intended-edit structure, unexpected alleles, junction architecture, and structural-event profiles across samples rather than interpreting each sample in isolation.

6. Deliver an evidence-weighted characterization report

Results are organized into observed events, supporting reads, allele or haplotype context, uncertainty, and recommended orthogonal follow-up. We clearly separate direct sequence evidence from interpretation and avoid labeling an event as biologically significant solely because it is detected.

Bioinformatics Analysis and Evidence Reporting

Analysis moduleTypical outputInterpretive purpose
Read QC and sample identityRead-length, quality, coverage, mapping, barcode/sample checksConfirms the dataset is suitable for the planned structural and allele-level analysis.
Target-region alignmentReference-aligned reads across the intended edit and flanksVisualizes the complete local allele rather than only the immediate edit window.
Small-variant callingSNVs and short indels with read supportConfirms intended base-level edits and identifies local bystander or secondary sequence changes.
Structural-variant analysisDeletions, insertions, inversions, duplications, translocations or complex rearrangementsDefines the larger editing outcomes that may be missed by short targeted assays.
Donor/vector integration analysisHost-donor junctions, insert composition, orientation and copy architectureDetermines whether intended or unintended exogenous sequence is incorporated at the locus.
Haplotype phasingEdit-to-variant linkage across long moleculesDistinguishes allele-specific outcomes and separates compound editing events.
Nominated off-target analysisPer-locus allele structures and event supportConfirms and characterizes edits at candidate sites supplied by the research program.
Genome-wide structural reviewSV calls, breakpoint evidence, copy-number context and prioritizationSupports broader genome-integrity studies when targeted analysis is not sufficient.
Optional methylation-aware analysisNative DNA methylation calls in the edited region and flanksAdds epigenetic context without requiring a separate chemical conversion assay.

Bioinformatics parameters are matched to the assay type and expected allele spectrum. Long-read gene-editing data are especially sensitive to reference choice, repetitive sequence, split alignment, PCR artifact, molecular barcodes strategy when used, and the treatment of low-frequency structural events. We therefore report supporting evidence rather than relying on a single automated variant table.

Characterization in a Gene-Therapy Development Context

For human genome-editing products, regulatory guidance increasingly emphasizes both off-target editing and loss of genome integrity. The FDA's January 2024 final guidance, Human Gene Therapy Products Incorporating Human Genome Editing, provides recommendations spanning product design, manufacturing/testing, nonclinical assessment, and clinical development.

In April 2026, FDA issued the draft guidance Safety Assessment of Genome Editing in Human Gene Therapy Products Using Next-Generation Sequencing. The draft is explicitly not for implementation and describes current recommendations for NGS-based nonclinical evaluation of off-target editing and loss of genome integrity.

Our service can generate sequencing evidence that may be useful within a sponsor's broader characterization program, but it does not by itself establish IND/BLA suitability, validated regulatory compliance, GMP release status, or an accepted off-target assessment strategy. The sponsor remains responsible for defining the product-specific risk framework, nomination methods, confirmation strategy, controls, acceptance criteria, validation status, and regulatory interpretation.

Sample Requirements and Project Entry Modes

Entry modeTypical materials/informationDesign notes
Targeted on-target characterizationEdited genomic DNA or cells plus unedited control; target coordinates; expected edit; guide/editor informationHigh-molecular-weight DNA is preferred when the expected event or flanking context is long. Exact input and QC are set after target review.
Knock-in / donor integration characterizationEdited material, donor/vector reference sequence, intended junctions, parental control when availableFull donor and backbone sequences improve interpretation of partial, tandem, inverted, or unexpected integrations.
Nominated off-target confirmationEdited and control material plus candidate off-target coordinates from sponsor or nomination assayTarget list and flanking span are selected to capture both small changes and plausible structural outcomes.
Long-read whole-genome characterizationHigh-quality high-molecular-weight genomic DNA from edited and matched control materialUsed when the research question includes large rearrangements, unknown integrations, genome-wide SVs, or allele phasing beyond defined loci.
Existing long-read data analysisFASTQ/BAM plus reference genome, target design, donor/vector sequences, sample metadataAnalysis-only projects can be accepted when the available data are technically appropriate for the intended characterization question.

Because gene-editing projects vary widely in cell type, genome, target size, mosaicism, donor architecture, and assay sensitivity requirements, we do not impose one universal input specification on this Solution page. Detailed submission requirements are finalized after project design; general handling information is available in our sample submission guideline.

Why CD Genomics for Gene-Editing Characterization?

PacBio and ONT are selected by question, not by catalog

Some projects need highly accurate local sequence. Others need maximum molecule length, native DNA signal, or broad whole-genome structural context. Our dual-platform model lets us select the technology around the edit architecture instead of forcing every sample into one instrument workflow.

We connect local alleles to genome-scale structure

A project can begin with a single target and expand if the data reveal a broader question. Targeted long-read sequencing, structural-variant analysis, long-read WGS, and optional methylation analysis can be connected within one interpretation framework rather than delivered as unrelated datasets.

We design around controls and false-negative risk

A structurally altered allele can be missed when PCR primers no longer bind, when a deletion extends outside the amplicon, or when repetitive sequence causes ambiguous mapping. We explicitly consider allelic dropout and assay boundary effects during design and recommend an orthogonal route when the first assay cannot test the relevant failure mode.

We keep research evidence separate from regulatory conclusions

Our reports describe what the reads support, what remains uncertain, and what may require orthogonal confirmation. We do not convert a sequencing observation into a clinical or regulatory claim without evidence outside the scope of this research-use service.

Independent Published Example: Long Reads Reveal Structural Outcomes Beyond Small Indels

Höijer I, Emmanouilidou A, Östlund R, et al. CRISPR-Cas9 induces large structural variants at on-target and off-target sites in vivo that segregate across generations. Nature Communications. 2022;13:627. doi:10.1038/s41467-022-28244-5. This open-access article is licensed under CC BY 4.0.

Background

The study asked whether standard views of CRISPR-Cas9 editing underestimate larger structural outcomes at both intended and off-target loci. The authors edited zebrafish with four guide RNAs and followed editing outcomes across more than 1,100 larvae, juvenile, and adult fish over two generations.

Methods

For detailed target-site characterization, the researchers generated long amplicons spanning 2.6–7.7 kb around Cas9 cleavage sites and sequenced them on the PacBio Sequel platform. They analyzed edited samples together with uninjected controls and used long reads to quantify and classify individual insertion and deletion alleles. Two F1 individuals were also examined by nanopore whole-genome sequencing to investigate apparent allelic imbalance that could not be resolved confidently by amplicon data alone.

Results

In 595 founder larvae analyzed across 20 pools, on-target editing outcomes ranged from a 4.8 kb deletion to a 1.4 kb insertion. Seven percent of on-target outcomes were structural variants of at least 50 bp; the corresponding fraction at confirmed off-target sites was 4%, giving a combined estimate of 6% across on- and off-target sites in founder larvae. The study also identified a 903 bp deletion at an off-target site that removed an entire exon of the unintended gene ywhaqb.

published Figure 5 showing size distribution of CRISPR induced on-target and off-target mutationsFigure 5 from Höijer et al. (2022), Nature Communications, CC BY 4.0: size distribution of CRISPR-Cas9-induced mutations at on-target and off-target sites, including a 903 bp off-target deletion spanning an exon.

Conclusion

This independent published example demonstrates why editing efficiency measured near a cut site is not equivalent to complete edit characterization. Long reads revealed a continuous spectrum from small indels to kilobase-scale structural variants and highlighted a second design lesson: targeted amplicon sequencing can itself suffer allelic dropout, so broader amplification-free or whole-genome analysis may be needed when the observed allele balance is biologically implausible. These findings support a tiered strategy in which the assay scale is expanded when the failure mode extends beyond the original target window.

FAQs

Sample Deliverables

1. On-target allele architecture report — intended edit, local SNVs/indels, large deletions/insertions, inversions, complex alleles, and read-level support summarized against the expected design.

2. Knock-in and exogenous sequence map — donor/vector junctions, inserted-fragment composition, orientation, tandem structure, partial integrations, and surrounding host sequence.

3. Haplotype and allele-phasing view — linkage between the intended edit and nearby heterozygous variants to distinguish allele-specific outcomes.

4. Nominated off-target characterization table — per-locus editing events, structural outcomes, supporting reads, and comparison with matched control material.

5. Optional genome-integrity report — long-read WGS structural variants, breakpoint evidence, copy-number context, unexpected integrations, and prioritized events for orthogonal follow-up.

sample gene editing characterization deliverables with allele maps structural variants and phasing

illustrative long-read gene editing allele map with deletion insertion inversion and donor integration

References

  1. Park SH, Cao M, Pan Y, et al. Comprehensive analysis and accurate quantification of unintended large gene modifications induced by CRISPR-Cas9 gene editing. Science Advances. 2022;8(42):eabo7676. doi:10.1126/sciadv.abo7676.
  2. Kosicki M, Allen F, Steward F, et al. Cas9-induced large deletions and small indels are controlled in a convergent fashion. Nature Communications. 2022;13:3422. doi:10.1038/s41467-022-30480-8.
  3. Tao J, Bauer DE, Chiarle R. Assessing and advancing the safety of CRISPR-Cas tools: from DNA to RNA editing. Nature Communications. 2023;14:212. doi:10.1038/s41467-023-35886-6.

For Research Use Only. Not for use in diagnostic or clinical procedures.

Get Your Instant Quote