Long-Read Sequencing for Plasmid and Production Cell Line Verification

Long-Read Sequencing for Plasmid and Production Cell Line Verification

long-read plasmid and production cell line verification for biopharma research

A plasmid can match its expected map at selected checkpoints and still contain an unexpected sequence change elsewhere. A production clone can express the intended product and still carry a complex integration locus composed of truncated, inverted, duplicated, or rearranged vector fragments. For biopharma and biotech research, these hidden structures can complicate clone selection, comparability studies, stability investigations, and interpretation of downstream expression data.

CD Genomics provides a long-read sequencing solution that connects whole-plasmid verification with production cell line genetic characterization. Using complementary PacBio HiFi and Oxford Nanopore Technologies (ONT) workflows, we help research teams confirm complete construct sequence, resolve host-vector and vector-vector junctions, reconstruct transgene integration architecture, assess clone-to-clone structural differences, and add native DNA methylation context when it is relevant to the research question.

Solution highlights

Discuss Your Verification Project

Why Long Reads Matter from Plasmid to Production Clone

Plasmid verification and production cell line characterization are often treated as separate analytical tasks, but they are part of the same molecular continuity question: what sequence was designed, what sequence was introduced, and what structure is actually present in the selected cell line? Conventional assays answer parts of this question well. Sanger sequencing is highly useful for confirming defined regions. PCR can verify expected junctions. qPCR or ddPCR can support targeted copy-number measurements. Short-read sequencing can screen the genome broadly. However, these approaches may require prior knowledge of the expected event and can struggle when several rearrangements occur on the same molecule or when repeated vector copies create ambiguous assemblies.

Long-read sequencing changes the unit of evidence from isolated fragments to continuous molecules. A read can span a plasmid backbone, a host-vector boundary, multiple vector fragments, or a long rearranged region. This makes it possible to reconstruct not only where an expression construct integrated, but also how the integrated copies are arranged. Published studies in recombinant CHO cells have shown that manufacturing cell lines can contain near-full-length tandem copies in one clone and highly fragmented, differently oriented vector copies in another, illustrating why integration-site identification alone may not describe the complete transgene locus.

We therefore use long reads as a structural characterization layer rather than as a blanket replacement for orthogonal methods. If the project requires absolute copy-number quantitation, a targeted qPCR or ddPCR assay can remain appropriate. If a single known base must be confirmed, Sanger sequencing may be sufficient. Long-read sequencing becomes especially valuable when the uncertainty concerns whole-construct identity, unexpected plasmid variants, long-range integration structure, concatemer architecture, repetitive sequence, complex rearrangements, or the relationship between sequence structure and epigenetic state.

What We Can Verify Across the Construct-to-Clone Continuum

Research questionLong-read evidenceTypical interpretation
Does the plasmid match the intended design?Full-length plasmid consensus, SNV/indel review, backbone and insert structureConfirms expected sequence or localizes unexpected sequence differences across the complete construct.
Is the plasmid population structurally uniform?Read-level comparison, coverage profile, alternative molecular structuresFlags evidence consistent with mixed plasmid structures, rearrangements, or subpopulations for follow-up.
Where did the construct integrate?Host-vector junction-spanning reads and genomic coordinatesIdentifies candidate integration locus or loci and surrounding host-genome context.
What is the integrated architecture?Vector-vector junctions, orientation, truncations, inversions, tandem or complex concatemer structureReconstructs the physical arrangement of integrated vector fragments rather than reporting only a breakpoint.
Did integration alter nearby genomic structure?Long-range alignments, SV calls, local assembly and breakpoint analysisSupports evaluation of local deletions, duplications, inversions, or other structural changes near the integration locus.
Are two clones or banks structurally consistent?Shared junctions, integrated sequence architecture, structural fingerprint comparisonProvides molecular evidence for clone comparison, bank comparison, or passage-related follow-up.
Is epigenetic context relevant to expression stability?Native ONT signal with methylation-aware analysis when plannedAdds methylation context to the same long-range molecules used for integration characterization.

The exact evidence package is selected around the development question. We do not treat every project as a fixed panel. A plasmid-only verification project may require a simple whole-construct workflow, whereas a production cell line investigation may combine targeted enrichment, long-read whole-genome data, local assembly, structural-variant analysis, and orthogonal confirmation of selected findings.

verification scope from plasmid map to integrated transgene architecture in a production cell line

Whole-Plasmid Verification Before Cell Line Development

A plasmid is more than the coding sequence. Promoters, enhancers, origins of replication, selection markers, regulatory elements, repetitive regions, linkers, and backbone sequences can all affect how a construct behaves or how confidently it can be traced through later development steps. Verifying only the insert or a limited series of amplicons can leave sequence outside those regions unexamined.

Our Full-Length Plasmid Sequencing service provides the core entry point for complete construct sequence verification. Depending on the project, analysis can compare the long-read consensus against the expected reference, report sequence differences, inspect coverage consistency, and evaluate whether read-level evidence supports a single dominant plasmid structure. This is useful before transfection, after cloning or plasmid propagation, when a construct has undergone repeated engineering cycles, or when the expected plasmid map no longer explains experimental behavior.

Questions a whole-plasmid workflow can address

Whole-plasmid long-read sequencing is not a substitute for every release or identity assay. Its value is comprehensive molecular context. The output can be used to decide which changes are biologically meaningful, which sites merit orthogonal confirmation, and which construct should move forward into cell line development research.

Production Cell Line Verification after Transgene Integration

Once a plasmid or expression construct is introduced into a host cell, the resulting genomic structure can be very different from the original circular map. Random integration can generate tandem copies, partial copies, inversions, head-to-head or tail-to-tail fusions, vector fragments, host-genome rearrangements, and unexpected junctions. Two clones expressing the same recombinant protein can therefore carry very different genetic architectures.

For production cell line research, long-read sequencing can connect host and vector sequence across the same molecule. This supports integration-site mapping and, critically, reconstruction of the structure surrounding those sites. The approach is particularly useful for CHO and other recombinant cell systems where concatemer complexity makes short-fragment reconstruction difficult.

Integration-site and junction characterization

Host-vector junction reads localize integration breakpoints and define the genomic context around the construct. When targeted enrichment is appropriate, long molecules extending from a known construct sequence into previously unknown host sequence can increase the efficiency of locus reconstruction. For genome-wide or multi-locus questions, broader long-read strategies may be preferable.

Concatemer and integrated-vector architecture

Integration architecture is more informative than a simple vector copy count. We can analyze the order and orientation of vector fragments, distinguish full-length from truncated copies, identify vector-vector fusion junctions, and reconstruct locally complex structures. Published CHO work has demonstrated cases with multiple nearly intact vector copies and other cases with highly fragmented copies in mixed orientations, underscoring why copy number alone may not describe the genomic state of the production clone.

Local host-genome rearrangements

Integration can coincide with deletions, duplications, inversions, or other structural changes in the host genome. Our Long-Read Variant Calling workflows can be incorporated when the project requires broader assessment of structural variation around the transgene locus or across the genome.

Optional methylation-aware characterization

For projects investigating expression instability or clone-to-clone regulatory differences, ONT native DNA sequencing can provide sequence and methylation information from the same raw signal. Our Long-Read Sequencing of DNA Methylation capability can be integrated when epigenetic context is part of the study hypothesis. Methylation measurements are interpreted as research evidence rather than as a standalone explanation for production phenotype.

Choosing PacBio HiFi, ONT, or a Combined Strategy

PacBio and ONT are complementary rather than interchangeable. Platform selection is driven by the structure that must be resolved, the desired sequence accuracy, the need for native DNA information, and whether the project is plasmid-focused, genome-wide, or targeted to a known transgene.

Project needPacBio HiFiOxford NanoporeSelection logic
Whole-plasmid consensus verificationStrong fit when high-consensus sequence accuracy is centralStrong fit with complete plasmid-spanning molecules and consensus analysisSelect according to construct size, accuracy requirements, throughput, and project scale.
Very long integration structures and concatemersUseful for accurate long-read reconstructionParticularly useful when very long native molecules help span complex architecturePrioritize physical span when multiple vector fragments or long host flanks must be connected.
Targeted integration-locus analysisProject dependentStrong fit for amplification-free Cas9-targeted strategies when appropriateUse enrichment when a focused locus can answer the question more efficiently than whole-genome sequencing.
Native DNA methylation contextProject dependentStrong fitUse ONT when sequence structure and native methylation are intentionally analyzed together.
High-confidence small-variant confirmation within long moleculesStrong fitPossible with appropriate consensus and orthogonal confirmation strategyMatch the platform and confirmation plan to the consequence of a small sequence difference.
Broad genome structural characterizationStrong fitStrong fitChoose based on required read length, genome complexity, project scale, and downstream analyses.

Our PacBio SMRT Sequencing Technology and Oxford Nanopore Sequencing Technology pages provide additional platform background. For this solution, however, the platform is a means to resolve the construct-to-clone question; it is not the starting point for study design.

Integrated Plasmid and Cell Line Verification Workflow

  1. Define the verification question — We review the construct map, host cell system, clone history, expected integration model, previous PCR/Sanger/TLA/WGS results, and the decision the study must support.
  2. Assess input material — Purified plasmid DNA and/or high-molecular-weight genomic DNA are evaluated for suitability. Existing sequencing data can also be reviewed as a starting point.
  3. Select sequencing and enrichment strategy — PacBio HiFi, ONT native genomic DNA, targeted enrichment, or a combined design is selected according to required accuracy and molecular span.
  4. Generate long-read data with assay-specific QC — Libraries are sequenced with attention to read quality, molecule-length distribution, coverage, mapping behavior, and target enrichment where applicable.
  5. Reconstruct plasmid and integration architecture — Reads are compared with the expected construct, host-vector and vector-vector junctions are identified, and local structures are assembled or phased.
  6. Compare clones, banks, or passages and report evidence — Structural features are summarized, discrepancies are prioritized for orthogonal confirmation, and a research report links sequence findings to the original verification question.

horizontal plasmid and production cell line verification workflow from construct map to structural report

Bioinformatics for Construct Identity and Integration Architecture

The central analytical challenge is not simply read alignment. A useful verification report must distinguish expected construct sequence from unexpected variation, separate host-vector from vector-vector junctions, and describe long-range architecture in a form that a cell line development team can interpret.

Analysis layerTypical outputsResearch value
Read and library QCRead-length distribution, quality summaries, mapping rate, coverage profileEstablishes whether the data support the planned plasmid or integration analysis.
Plasmid consensus analysisReference-aligned consensus, sequence differences, per-region supportConfirms complete construct identity and localizes discrepancies.
Read-level structural inspectionAlternative structures, breakpoint-supporting reads, coverage irregularitiesHelps distinguish a simple consensus difference from evidence of structural heterogeneity.
Host-vector junction mappingGenomic breakpoint coordinates and junction-spanning moleculesIdentifies candidate integration loci and host sequence context.
Vector-vector junction analysisOrientation, adjacency, truncation, inversion, head-to-head/tail-to-tail relationshipsReconstructs concatemer or fragmented transgene architecture.
Local assembly and structural variant analysisIntegration-locus reconstruction, local host rearrangements, phased structureConnects several breakpoints into one interpretable genomic model.
Copy-architecture assessmentLong-read evidence supporting number and organization of vector fragmentsProvides structural copy context; absolute copy-number assays can be added orthogonally if required.
Optional methylation analysisNative DNA methylation profiles across transgene and flanking host regionsAdds epigenetic context for research on expression state or stability.
Comparative clone/bank analysisShared and discordant junctions, architecture comparison, structural fingerprintSupports clone selection, bank comparison, or passage/stability investigations.

Projects requiring custom analysis can also enter through our Long-Read Sequencing Data Analysis Services, including projects where sequencing has already been completed elsewhere but the integration structure remains unresolved.

Characterization Context for Biopharma Development Research

Expression-construct and cell-substrate characterization are longstanding concerns in biotechnology development. ICH Q5B discusses the need to understand the detailed expression construct, including an annotated plasmid sequence and inserted coding regions with flanking junctions, and it also addresses characterization of the expression construct in cell banks for features such as copy number, insertions or deletions, and integration sites. ICH Q5D addresses the derivation and characterization of cell substrates used for biotechnology products.

These frameworks explain why construct identity and production-cell genetic architecture matter, but they do not mean that any single sequencing workflow automatically satisfies a regulatory requirement. Our service is designed as research-use sequencing and characterization support. Project-specific regulatory strategy, validated release testing, acceptance criteria, GMP requirements, and filing decisions remain the responsibility of the sponsor and its quality/regulatory teams.

Research safeguards built into the solution

  • Reference traceability: analysis starts from the intended construct map and sequence version supplied for the project.
  • Unexpected-event reporting: we distinguish expected design features from sequence or structural findings that merit follow-up.
  • Orthogonal confirmation strategy: high-consequence findings can be prioritized for Sanger, PCR, qPCR/ddPCR, Southern blot, or another independent assay as appropriate.
  • Architecture before interpretation: copy number is not treated as a substitute for knowing copy orientation, truncation, junctions, and local genomic context.
  • No automatic phenotype claims: sequence or methylation differences are not assumed to explain productivity or stability without supporting experimental evidence.

Sample and Project Entry Modes

Because this is a solution rather than a fixed single assay, input requirements depend on whether the project begins before or after cell line generation. We confirm final material requirements after reviewing the construct, host cell system, desired coverage or enrichment strategy, and whether native high-molecular-weight DNA must be preserved.

Entry modeTypical inputWhat we can plan
Plasmid-only verificationPurified plasmid DNA plus expected sequence/mapWhole-plasmid consensus, variant review, structural consistency assessment.
Plasmid + candidate clonesReference plasmid plus high-quality genomic DNA from selected clonesConstruct verification followed by integration-site and architecture comparison.
Cell line / cell bank characterizationHigh-molecular-weight genomic DNA from clone, MCB/WCB research samples, or defined passage pointsLong-range integration analysis, local structural variation, comparative structural fingerprinting.
Targeted follow-upGenomic DNA plus known vector sequence and prior breakpoint evidenceTargeted enrichment and focused reconstruction of a difficult integration locus.
Existing dataPacBio/ONT reads, assemblies, BAM/CRAM, prior TLA/WGS/PCR resultsReanalysis, local assembly, junction reconstruction, or cross-platform evidence integration.

For native ONT workflows, DNA extraction and handling should preserve long molecules because molecule length directly affects the ability to span long integration structures. For plasmid verification, purity and a clearly versioned reference sequence are central to interpretation. We avoid publishing one universal input threshold for all solution modes because targeted enrichment, whole-genome sequencing, plasmid sequencing, and methylation-aware analysis have different material requirements.

Why Choose CD Genomics for Construct-to-Clone Verification

One solution connects the original plasmid to the derived cell line

Instead of treating plasmid sequence and cell line integration as unrelated tests, we design the project around molecular continuity. The same construct reference can be carried from whole-plasmid verification into host-vector junction analysis and final integration-architecture reporting.

PacBio and ONT can be selected around the question

Our CDL2 long-read platform is designed around complementary PacBio and ONT capabilities. High-consensus sequence accuracy, very long native molecules, targeted enrichment, and methylation-aware analysis can be combined or separated according to what the project actually needs.

Structural interpretation goes beyond a breakpoint list

For complex production clones, reporting one integration coordinate may be insufficient. We focus on interpretable architecture: which vector fragments are present, how they are oriented, how they connect to one another and to the host genome, and what local rearrangements are supported by the reads.

Custom bioinformatics supports unusual constructs and cell systems

Non-standard vectors, multiple transgenes, complex cassettes, repeat-rich constructs, non-CHO host systems, and projects with prior mixed-platform data can be handled as custom research workflows rather than forced into a rigid catalog assay.

FAQs

Sample Deliverables

1. Whole-plasmid consensus report — complete construct sequence aligned to the expected reference, with sequence differences and coverage support summarized for review.

2. Integration architecture map — host-vector and vector-vector junctions displayed with copy orientation, truncation, inversion, concatemer relationships, and local genomic context.

3. Clone or bank comparison matrix — shared and discordant junctions and structural features across selected clones, cell-bank samples, or passage points.

4. Optional methylation-aware locus view — native ONT methylation signal overlaid on transgene and flanking host sequence when included in the study design.

5. Evidence-based interpretation report — prioritized discrepancies, recommended orthogonal follow-up, and a clear distinction between observed sequence structure and biological interpretation.

sample deliverables for plasmid consensus integration architecture and clone comparison

illustrative long-read host vector integration architecture with complex concatemer junctions

References

  1. Brown SD, Dreolini L, Wilson JF, Balasundaram M, Holt RA. Complete sequence verification of plasmid DNA using the Oxford Nanopore Technologies' MinION device. BMC Bioinformatics. 2023;24:116. doi:10.1186/s12859-023-05226-y.
  2. Clappier C, Böttner D, Heinzelmann D, Stadermann A, Schulz P, Schmidt M, Lindner B. Deciphering integration loci of CHO manufacturing cell lines using long read nanopore sequencing. New Biotechnology. 2023;75:31-39. doi:10.1016/j.nbt.2023.03.003.
  3. Leitner K, Motheramgari K, Borth N, Marx N. Nanopore Cas9-targeted sequencing enables accurate and simultaneous identification of transgene integration sites, their structure and epigenetic status in recombinant Chinese hamster ovary cells. Biotechnology and Bioengineering. 2023;120(9):2403-2418. doi:10.1002/bit.28382.
  4. Slesarev A, Viswanathan L, Tang Y, et al. CRISPR/Cas9 targeted CAPTURE of mammalian genomic regions for characterization by NGS. Scientific Reports. 2019;9:3587. doi:10.1038/s41598-019-39667-4.

For Research Use Only. Not for use in diagnostic or clinical procedures.

Get Your Instant Quote