
A plasmid can match its expected map at selected checkpoints and still contain an unexpected sequence change elsewhere. A production clone can express the intended product and still carry a complex integration locus composed of truncated, inverted, duplicated, or rearranged vector fragments. For biopharma and biotech research, these hidden structures can complicate clone selection, comparability studies, stability investigations, and interpretation of downstream expression data.
CD Genomics provides a long-read sequencing solution that connects whole-plasmid verification with production cell line genetic characterization. Using complementary PacBio HiFi and Oxford Nanopore Technologies (ONT) workflows, we help research teams confirm complete construct sequence, resolve host-vector and vector-vector junctions, reconstruct transgene integration architecture, assess clone-to-clone structural differences, and add native DNA methylation context when it is relevant to the research question.
Solution highlights
Plasmid verification and production cell line characterization are often treated as separate analytical tasks, but they are part of the same molecular continuity question: what sequence was designed, what sequence was introduced, and what structure is actually present in the selected cell line? Conventional assays answer parts of this question well. Sanger sequencing is highly useful for confirming defined regions. PCR can verify expected junctions. qPCR or ddPCR can support targeted copy-number measurements. Short-read sequencing can screen the genome broadly. However, these approaches may require prior knowledge of the expected event and can struggle when several rearrangements occur on the same molecule or when repeated vector copies create ambiguous assemblies.
Long-read sequencing changes the unit of evidence from isolated fragments to continuous molecules. A read can span a plasmid backbone, a host-vector boundary, multiple vector fragments, or a long rearranged region. This makes it possible to reconstruct not only where an expression construct integrated, but also how the integrated copies are arranged. Published studies in recombinant CHO cells have shown that manufacturing cell lines can contain near-full-length tandem copies in one clone and highly fragmented, differently oriented vector copies in another, illustrating why integration-site identification alone may not describe the complete transgene locus.
We therefore use long reads as a structural characterization layer rather than as a blanket replacement for orthogonal methods. If the project requires absolute copy-number quantitation, a targeted qPCR or ddPCR assay can remain appropriate. If a single known base must be confirmed, Sanger sequencing may be sufficient. Long-read sequencing becomes especially valuable when the uncertainty concerns whole-construct identity, unexpected plasmid variants, long-range integration structure, concatemer architecture, repetitive sequence, complex rearrangements, or the relationship between sequence structure and epigenetic state.
| Research question | Long-read evidence | Typical interpretation |
| Does the plasmid match the intended design? | Full-length plasmid consensus, SNV/indel review, backbone and insert structure | Confirms expected sequence or localizes unexpected sequence differences across the complete construct. |
| Is the plasmid population structurally uniform? | Read-level comparison, coverage profile, alternative molecular structures | Flags evidence consistent with mixed plasmid structures, rearrangements, or subpopulations for follow-up. |
| Where did the construct integrate? | Host-vector junction-spanning reads and genomic coordinates | Identifies candidate integration locus or loci and surrounding host-genome context. |
| What is the integrated architecture? | Vector-vector junctions, orientation, truncations, inversions, tandem or complex concatemer structure | Reconstructs the physical arrangement of integrated vector fragments rather than reporting only a breakpoint. |
| Did integration alter nearby genomic structure? | Long-range alignments, SV calls, local assembly and breakpoint analysis | Supports evaluation of local deletions, duplications, inversions, or other structural changes near the integration locus. |
| Are two clones or banks structurally consistent? | Shared junctions, integrated sequence architecture, structural fingerprint comparison | Provides molecular evidence for clone comparison, bank comparison, or passage-related follow-up. |
| Is epigenetic context relevant to expression stability? | Native ONT signal with methylation-aware analysis when planned | Adds methylation context to the same long-range molecules used for integration characterization. |
The exact evidence package is selected around the development question. We do not treat every project as a fixed panel. A plasmid-only verification project may require a simple whole-construct workflow, whereas a production cell line investigation may combine targeted enrichment, long-read whole-genome data, local assembly, structural-variant analysis, and orthogonal confirmation of selected findings.

A plasmid is more than the coding sequence. Promoters, enhancers, origins of replication, selection markers, regulatory elements, repetitive regions, linkers, and backbone sequences can all affect how a construct behaves or how confidently it can be traced through later development steps. Verifying only the insert or a limited series of amplicons can leave sequence outside those regions unexamined.
Our Full-Length Plasmid Sequencing service provides the core entry point for complete construct sequence verification. Depending on the project, analysis can compare the long-read consensus against the expected reference, report sequence differences, inspect coverage consistency, and evaluate whether read-level evidence supports a single dominant plasmid structure. This is useful before transfection, after cloning or plasmid propagation, when a construct has undergone repeated engineering cycles, or when the expected plasmid map no longer explains experimental behavior.
Whole-plasmid long-read sequencing is not a substitute for every release or identity assay. Its value is comprehensive molecular context. The output can be used to decide which changes are biologically meaningful, which sites merit orthogonal confirmation, and which construct should move forward into cell line development research.
Once a plasmid or expression construct is introduced into a host cell, the resulting genomic structure can be very different from the original circular map. Random integration can generate tandem copies, partial copies, inversions, head-to-head or tail-to-tail fusions, vector fragments, host-genome rearrangements, and unexpected junctions. Two clones expressing the same recombinant protein can therefore carry very different genetic architectures.
For production cell line research, long-read sequencing can connect host and vector sequence across the same molecule. This supports integration-site mapping and, critically, reconstruction of the structure surrounding those sites. The approach is particularly useful for CHO and other recombinant cell systems where concatemer complexity makes short-fragment reconstruction difficult.
Host-vector junction reads localize integration breakpoints and define the genomic context around the construct. When targeted enrichment is appropriate, long molecules extending from a known construct sequence into previously unknown host sequence can increase the efficiency of locus reconstruction. For genome-wide or multi-locus questions, broader long-read strategies may be preferable.
Integration architecture is more informative than a simple vector copy count. We can analyze the order and orientation of vector fragments, distinguish full-length from truncated copies, identify vector-vector fusion junctions, and reconstruct locally complex structures. Published CHO work has demonstrated cases with multiple nearly intact vector copies and other cases with highly fragmented copies in mixed orientations, underscoring why copy number alone may not describe the genomic state of the production clone.
Integration can coincide with deletions, duplications, inversions, or other structural changes in the host genome. Our Long-Read Variant Calling workflows can be incorporated when the project requires broader assessment of structural variation around the transgene locus or across the genome.
For projects investigating expression instability or clone-to-clone regulatory differences, ONT native DNA sequencing can provide sequence and methylation information from the same raw signal. Our Long-Read Sequencing of DNA Methylation capability can be integrated when epigenetic context is part of the study hypothesis. Methylation measurements are interpreted as research evidence rather than as a standalone explanation for production phenotype.
PacBio and ONT are complementary rather than interchangeable. Platform selection is driven by the structure that must be resolved, the desired sequence accuracy, the need for native DNA information, and whether the project is plasmid-focused, genome-wide, or targeted to a known transgene.
| Project need | PacBio HiFi | Oxford Nanopore | Selection logic |
| Whole-plasmid consensus verification | Strong fit when high-consensus sequence accuracy is central | Strong fit with complete plasmid-spanning molecules and consensus analysis | Select according to construct size, accuracy requirements, throughput, and project scale. |
| Very long integration structures and concatemers | Useful for accurate long-read reconstruction | Particularly useful when very long native molecules help span complex architecture | Prioritize physical span when multiple vector fragments or long host flanks must be connected. |
| Targeted integration-locus analysis | Project dependent | Strong fit for amplification-free Cas9-targeted strategies when appropriate | Use enrichment when a focused locus can answer the question more efficiently than whole-genome sequencing. |
| Native DNA methylation context | Project dependent | Strong fit | Use ONT when sequence structure and native methylation are intentionally analyzed together. |
| High-confidence small-variant confirmation within long molecules | Strong fit | Possible with appropriate consensus and orthogonal confirmation strategy | Match the platform and confirmation plan to the consequence of a small sequence difference. |
| Broad genome structural characterization | Strong fit | Strong fit | Choose based on required read length, genome complexity, project scale, and downstream analyses. |
Our PacBio SMRT Sequencing Technology and Oxford Nanopore Sequencing Technology pages provide additional platform background. For this solution, however, the platform is a means to resolve the construct-to-clone question; it is not the starting point for study design.

The central analytical challenge is not simply read alignment. A useful verification report must distinguish expected construct sequence from unexpected variation, separate host-vector from vector-vector junctions, and describe long-range architecture in a form that a cell line development team can interpret.
| Analysis layer | Typical outputs | Research value |
| Read and library QC | Read-length distribution, quality summaries, mapping rate, coverage profile | Establishes whether the data support the planned plasmid or integration analysis. |
| Plasmid consensus analysis | Reference-aligned consensus, sequence differences, per-region support | Confirms complete construct identity and localizes discrepancies. |
| Read-level structural inspection | Alternative structures, breakpoint-supporting reads, coverage irregularities | Helps distinguish a simple consensus difference from evidence of structural heterogeneity. |
| Host-vector junction mapping | Genomic breakpoint coordinates and junction-spanning molecules | Identifies candidate integration loci and host sequence context. |
| Vector-vector junction analysis | Orientation, adjacency, truncation, inversion, head-to-head/tail-to-tail relationships | Reconstructs concatemer or fragmented transgene architecture. |
| Local assembly and structural variant analysis | Integration-locus reconstruction, local host rearrangements, phased structure | Connects several breakpoints into one interpretable genomic model. |
| Copy-architecture assessment | Long-read evidence supporting number and organization of vector fragments | Provides structural copy context; absolute copy-number assays can be added orthogonally if required. |
| Optional methylation analysis | Native DNA methylation profiles across transgene and flanking host regions | Adds epigenetic context for research on expression state or stability. |
| Comparative clone/bank analysis | Shared and discordant junctions, architecture comparison, structural fingerprint | Supports clone selection, bank comparison, or passage/stability investigations. |
Projects requiring custom analysis can also enter through our Long-Read Sequencing Data Analysis Services, including projects where sequencing has already been completed elsewhere but the integration structure remains unresolved.
Expression-construct and cell-substrate characterization are longstanding concerns in biotechnology development. ICH Q5B discusses the need to understand the detailed expression construct, including an annotated plasmid sequence and inserted coding regions with flanking junctions, and it also addresses characterization of the expression construct in cell banks for features such as copy number, insertions or deletions, and integration sites. ICH Q5D addresses the derivation and characterization of cell substrates used for biotechnology products.
These frameworks explain why construct identity and production-cell genetic architecture matter, but they do not mean that any single sequencing workflow automatically satisfies a regulatory requirement. Our service is designed as research-use sequencing and characterization support. Project-specific regulatory strategy, validated release testing, acceptance criteria, GMP requirements, and filing decisions remain the responsibility of the sponsor and its quality/regulatory teams.
Because this is a solution rather than a fixed single assay, input requirements depend on whether the project begins before or after cell line generation. We confirm final material requirements after reviewing the construct, host cell system, desired coverage or enrichment strategy, and whether native high-molecular-weight DNA must be preserved.
| Entry mode | Typical input | What we can plan |
| Plasmid-only verification | Purified plasmid DNA plus expected sequence/map | Whole-plasmid consensus, variant review, structural consistency assessment. |
| Plasmid + candidate clones | Reference plasmid plus high-quality genomic DNA from selected clones | Construct verification followed by integration-site and architecture comparison. |
| Cell line / cell bank characterization | High-molecular-weight genomic DNA from clone, MCB/WCB research samples, or defined passage points | Long-range integration analysis, local structural variation, comparative structural fingerprinting. |
| Targeted follow-up | Genomic DNA plus known vector sequence and prior breakpoint evidence | Targeted enrichment and focused reconstruction of a difficult integration locus. |
| Existing data | PacBio/ONT reads, assemblies, BAM/CRAM, prior TLA/WGS/PCR results | Reanalysis, local assembly, junction reconstruction, or cross-platform evidence integration. |
For native ONT workflows, DNA extraction and handling should preserve long molecules because molecule length directly affects the ability to span long integration structures. For plasmid verification, purity and a clearly versioned reference sequence are central to interpretation. We avoid publishing one universal input threshold for all solution modes because targeted enrichment, whole-genome sequencing, plasmid sequencing, and methylation-aware analysis have different material requirements.
Instead of treating plasmid sequence and cell line integration as unrelated tests, we design the project around molecular continuity. The same construct reference can be carried from whole-plasmid verification into host-vector junction analysis and final integration-architecture reporting.
Our CDL2 long-read platform is designed around complementary PacBio and ONT capabilities. High-consensus sequence accuracy, very long native molecules, targeted enrichment, and methylation-aware analysis can be combined or separated according to what the project actually needs.
For complex production clones, reporting one integration coordinate may be insufficient. We focus on interpretable architecture: which vector fragments are present, how they are oriented, how they connect to one another and to the host genome, and what local rearrangements are supported by the reads.
Non-standard vectors, multiple transgenes, complex cassettes, repeat-rich constructs, non-CHO host systems, and projects with prior mixed-platform data can be handled as custom research workflows rather than forced into a rigid catalog assay.
Not always. Sanger sequencing is excellent for confirming defined regions. Long-read whole-plasmid verification is most useful when you want continuous evidence across the entire construct, when many primer reactions would otherwise be required, or when unexpected structural changes outside the planned amplicons are a concern.
Long-read data can provide strong evidence about copy architecture and can support copy-number estimation, especially when reads resolve individual integrated fragments and junctions. If an absolute copy-number measurement is a critical project requirement, we generally recommend interpreting long-read structure together with an orthogonal quantitative method such as ddPCR or qPCR.
Yes, when read length and coverage adequately span the relevant junctions. Long-read alignments and local assembly can distinguish full-length from truncated vector fragments, identify inversion and orientation, and connect vector-vector and host-vector junctions into a structural model.
Targeted enrichment can be useful when the vector sequence is known and the main goal is to recover long molecules extending through vector-vector or vector-host junctions without sequencing the entire genome deeply. Whole-genome long-read sequencing may be preferable when integration sites are unknown, multiple loci are suspected, or broader host-genome structural characterization is required.
For appropriately designed ONT native DNA projects, methylation-aware analysis can be performed from the same raw signal used for sequence analysis. This can add epigenetic context around the transgene and flanking host sequence. Methylation findings should be interpreted alongside expression and stability data rather than assumed to establish causality on their own.
Yes, as a research characterization study. Shared junctions and long-range architecture can provide a structural fingerprint for comparing defined clones, bank samples, or passage points. The experimental design should be matched to the degree of change the study needs to detect and to any orthogonal assays already in use.
No. ICH Q5B and Q5D provide important context for expression-construct and cell-substrate characterization, but our service is research-use sequencing support. Regulatory compliance, validation, GMP status, acceptance criteria, and submission strategy must be established by the sponsor and its quality and regulatory teams.
1. Whole-plasmid consensus report — complete construct sequence aligned to the expected reference, with sequence differences and coverage support summarized for review.
2. Integration architecture map — host-vector and vector-vector junctions displayed with copy orientation, truncation, inversion, concatemer relationships, and local genomic context.
3. Clone or bank comparison matrix — shared and discordant junctions and structural features across selected clones, cell-bank samples, or passage points.
4. Optional methylation-aware locus view — native ONT methylation signal overlaid on transgene and flanking host sequence when included in the study design.
5. Evidence-based interpretation report — prioritized discrepancies, recommended orthogonal follow-up, and a clear distinction between observed sequence structure and biological interpretation.


References
For Research Use Only. Not for use in diagnostic or clinical procedures.