Long-Read Sequencing for Biomarker Discovery

Long-Read Sequencing for Biomarker Discovery

long-read sequencing biomarker discovery integrating structural variants isoforms haplotypes and methylation

Biomarker discovery often stalls after an initial association because the underlying molecular feature is more complex than a short-read signal suggests. Long-read sequencing expands the searchable space by resolving complete structural variants, phased haplotypes, full-length transcript isoforms, fusion transcripts, and native modification patterns that can be obscured by fragmented sequencing.

CD Genomics provides a biomarker discovery solution built around complementary PacBio HiFi and Oxford Nanopore Technologies (ONT) workflows. We help research, biotech, and biopharma teams move from an unresolved genomic or transcriptomic signal to a prioritized, validation-ready set of molecular candidates with sequence, structure, phase, expression, and modification evidence matched to the research question.

Solution highlights

Discuss Your Biomarker Project

Why Long Reads Expand the Biomarker Search Space

Many discovery programs begin with short-read WGS, RNA-seq, single-cell RNA-seq, genotyping arrays, or targeted panels. These technologies are efficient for broad screening, but some candidate regions remain unresolved because the biological feature is longer than the read, embedded in repeats, structurally rearranged, or distributed across multiple transcript exons. In those settings, deeper short-read sequencing may add coverage without adding the missing molecular context.

Long reads change the unit of evidence. A single molecule can span a structural-variant breakpoint, carry multiple heterozygous variants for phasing, cover an entire transcript isoform, or retain native modification information. This allows the discovery team to ask not only whether a locus is associated with a phenotype, but also which molecular configuration is present and whether that configuration is linked to a specific transcript, haplotype, or regulatory state.

We do not position long-read sequencing as a universal replacement for short-read discovery. Large cohorts focused on common SNVs or gene-level expression may still be more efficiently screened with high-throughput short-read methods. Long-read sequencing is most valuable when molecular structure, phase, isoform identity, repetitive context, or modification state changes the interpretation of the candidate signal.

Biomarker Classes Resolved by Long-Read Sequencing

Structural-Variant and Complex-Locus Candidates

Candidate biomarkers can be driven by insertions, deletions, inversions, duplications, translocations, repeat expansions, or compound rearrangements that are difficult to reconstruct from short fragments. Long reads can span breakpoints and repetitive flanking sequence, providing direct evidence for the architecture of the event. For projects centered on genome-wide discovery, our Human Whole Genome Sequencing workflow provides long-read coverage for genome-scale variant analysis. When the central question is a complex rearrangement, Human Genome Structural Variation Detection provides a focused route for SV discovery and interpretation.

Phased and Haplotype-Linked Candidates

A candidate may depend on whether variants occur in cis or trans, whether a regulatory change sits on the same molecule as a coding variant, or whether multiple variants define a risk- or response-associated haplotype. Long reads phase variants directly across longer genomic intervals, allowing candidate interpretation at the haplotype level rather than as isolated variant calls. This is particularly useful for highly polymorphic loci, duplicated regions, pharmacogenomic genes, and studies where allele-specific effects are biologically relevant.

Full-Length Isoform and Fusion-Transcript Candidates

Gene-level abundance can hide changes in transcript structure. A biologically important signal may be an isoform switch, novel splice junction, alternative transcription start or end site, intron-retaining transcript, or fusion transcript rather than a change in total gene expression. Our Full-Length Transcript Sequencing (Iso-Seq) service captures complete transcript structures for isoform discovery and comparison. Long-read transcriptomics can therefore turn an ambiguous gene-level association into a defined transcript-level candidate that can be carried into targeted validation.

Single-Cell Genotype and Isoform Candidates

Single-cell expression data are powerful for defining cell states, but short-read libraries can lose the full molecular structure needed to connect cell identity with isoforms, fusion genes, or other transcript-linked genetic features. Our Single-Cell Full-Length Transcriptome Sequencing workflow adds full-length transcript information to cell-resolved studies, supporting candidate discovery in heterogeneous systems where a signal may be restricted to a rare cell population or specific cellular state.

Methylation and Native-Modification Candidates

Long-read single-molecule sequencing can connect sequence context and epigenetic state on the same molecules. For DNA-focused studies, Long-Read Sequencing of DNA Methylation supports methylation-aware analysis alongside long-range sequence information. For RNA-focused projects where native transcript information and RNA modification patterns are central to the hypothesis, Nanopore Direct RNA Sequencing can preserve molecular features that are not retained in conventional cDNA-based workflows.

long-read biomarker classes structural variants haplotypes isoforms fusions methylation and single-cell features

Choosing PacBio HiFi, ONT, or an Integrated Strategy

The best long-read platform depends on the candidate biomarker class and the evidence required to prioritize it. We select the platform around the research question rather than forcing every project into the same workflow.

Research need PacBio HiFi Oxford Nanopore Typical decision logic
High-confidence genome variant discovery Strong fit Strong fit Choose according to locus complexity, required read span, and project design.
Complex SV and repeat-spanning analysis Strong fit Strong fit, especially when very long molecules add value Use the platform that best resolves the physical span of the event.
Haplotype-resolved analysis Strong fit Strong fit Long molecules allow linked variants to be interpreted together.
Full-length transcript isoforms Strong fit through Iso-Seq Strong fit through cDNA or direct RNA workflows Choose according to accuracy, native-RNA requirements, and isoform question.
Native DNA/RNA modification analysis Project dependent Particularly useful for native signal analysis Prefer native-molecule workflows when modification state is part of the candidate definition.
Rapid focused follow-up Project dependent Flexible for targeted and adaptive strategies Match throughput and enrichment strategy to the candidate set.

For some biomarker programs, the most efficient design is sequential: broad discovery identifies a difficult candidate class, long-read sequencing resolves the molecular structure, and a separate orthogonal assay validates the prioritized candidate in a larger cohort. For others, long-read WGS or full-length transcriptomics is the primary discovery layer from the beginning.

Integrated Biomarker Discovery Workflow

  1. Research question and evidence review — We review existing WGS, RNA-seq, single-cell, array, association, or pilot data and define the unresolved molecular question.
  2. Candidate class and sample assessment — We determine whether the project is primarily genomic, transcriptomic, epigenetic, single-cell, or multi-layer and evaluate material quality.
  3. Platform and library strategy — PacBio HiFi, ONT genomic DNA, full-length cDNA, direct RNA, or a combined strategy is selected according to the required molecular evidence.
  4. Long-read sequencing and assay-specific QC — Libraries are sequenced with QC focused on read quality, molecule length, alignment behavior, coverage distribution, and assay-specific metrics.
  5. Feature discovery and molecular resolution — We call and characterize SVs, phased variants, isoforms, fusions, methylation patterns, or other planned features using analysis pipelines matched to the data type.
  6. Candidate integration and prioritization — Molecular features are linked to phenotype, condition, cell state, pathway, or prior short-read evidence to produce a ranked shortlist for downstream validation.

The workflow image is intentionally presented as one horizontal decision path. Detailed project-specific steps remain in the accompanying text so the process can be adapted without implying a rigid one-size-fits-all protocol.

horizontal long-read biomarker discovery workflow from research question to candidate shortlist

Bioinformatics and Candidate Prioritization

A long-read experiment can generate thousands of structural, transcript, or modification features. Biomarker discovery therefore requires more than variant calling. Our analysis focuses on converting the raw feature space into candidates that can be evaluated biologically and technically.

Analysis layer Examples of outputs Biomarker-discovery value
Read and library QC Read-length distribution, mapping summaries, coverage profiles Confirms whether the dataset supports the intended feature class.
Genome variation SNVs/indels where relevant, SVs, CNVs, breakpoint structures Identifies sequence and structural candidates.
Haplotype analysis Phased variants, allele-specific structures, haplotype blocks Links variants into biologically interpretable molecular configurations.
Transcript structure Full-length isoforms, splice junctions, fusion transcripts, TSS/TES patterns Identifies transcript-level candidates hidden by gene-level counts.
Differential analysis Condition-associated isoform, variant, or modification patterns Prioritizes features associated with the experimental contrast.
Methylation/modification Site- or region-level modification profiles, haplotype-linked patterns Adds regulatory evidence to sequence-based candidates.
Single-cell integration Cell-type-specific isoforms, mutations, fusions, immune-receptor features Connects molecular candidates to cellular context.
Functional interpretation Gene annotation, pathway context, domain impact, known-locus overlap Helps rank candidates for validation rather than reporting an unfiltered list.

Candidate prioritization can incorporate recurrence across biological replicates, effect direction, molecular plausibility, orthogonal evidence, cell-type specificity, locus complexity, and compatibility with a downstream validation assay. We distinguish exploratory associations from stronger mechanistic evidence and do not present discovery-stage candidates as validated diagnostic markers.

Biopharma and Biotech Research Applications

Translational and Disease-Mechanism Research

Long-read sequencing can resolve molecular features that explain why a disease-associated locus or expression signal behaves differently across samples. Examples include a structural rearrangement that changes gene dosage, a phased haplotype linked to expression, a disease-linked transcript isoform, or a methylation state associated with a regulatory region. These features can provide a more precise research hypothesis for downstream experiments.

Drug-Response and Mechanism-of-Action Studies

Treatment-response studies often produce broad expression signatures. Long-read follow-up can ask whether the response is associated with a particular isoform, fusion, allele-specific transcript, structural variant, or epigenetic configuration. Candidate markers can then be prioritized for orthogonal testing in additional samples or experimental models.

Target and Pathway Discovery

A novel transcript isoform, fusion, or complex genomic rearrangement can alter protein domains, regulatory architecture, or pathway membership in ways that are missed by gene-level analysis. Long-read sequencing provides the full molecular sequence needed to predict open reading frames, annotate altered domains, and define candidate mechanisms for functional follow-up.

Pharmacogenomic and Highly Polymorphic Loci

Genes with segmental duplications, pseudogenes, repeats, or extensive haplotype diversity can be difficult to characterize with short reads. Long reads can improve molecular resolution across the locus and support phased candidate analysis. Research conclusions remain project-specific and should be validated with an appropriate independent method before broader use.

Cell-State and Immune Biomarker Research

In heterogeneous tissues, a candidate feature may be present only in a specific malignant, stromal, or immune cell population. Long-read single-cell workflows can connect full-length transcripts and genetic features to cell-state information, supporting research into lineage, clonal structure, immune receptor diversity, and cell-type-specific isoform usage.

From Discovery Signal to Validation-Ready Candidate

The most useful output of a biomarker discovery project is not the longest feature list; it is a smaller set of candidates with enough evidence to justify the next experiment. We therefore structure projects around evidence progression.

Discovery evidence may include a recurrent SV, condition-associated isoform, phased allele, or methylation pattern. Context evidence asks whether the feature is expressed in the relevant cell type, affects a plausible gene or pathway, or co-occurs with other molecular changes. Technical evidence evaluates read support, mapping ambiguity, breakpoint or isoform structure, and sample-level reproducibility. Validation planning identifies a practical orthogonal method such as targeted sequencing, PCR-based confirmation, digital PCR, targeted expression analysis, or another assay selected for the candidate class.

This separation is important in biopharma research. A statistically interesting long-read signal is not automatically a robust biomarker. The discovery page is designed to help teams create candidates that are structurally well defined and technically traceable, while leaving clinical qualification, diagnostic cutoff definition, and regulated validation outside the scope of this research-use service.

Sample and Project Entry Requirements

Biomarker discovery projects can enter at different stages. Final requirements depend on the selected long-read assay, so we confirm acceptance criteria during project design rather than applying one universal input specification.

Project entry Typical material Planning note
Genome-wide structural or haplotype discovery High-molecular-weight genomic DNA For current human long-read WGS workflows, ≥10 μg DNA and A260/280 of 1.8–2.0 are listed as service guidance; confirm final requirements before shipment.
Full-length transcript biomarker discovery High-integrity total RNA or poly(A)+ RNA Input and integrity targets depend on Iso-Seq or ONT RNA workflow and study design.
Single-cell follow-up Viable cells, prepared single-cell material, or compatible amplified cDNA depending on workflow Review library chemistry and barcode structure before project initiation.
Native methylation discovery High-quality genomic DNA suitable for native long-read sequencing Avoid amplification when native modification information is required.
Existing-data entry PacBio, ONT, short-read, single-cell, association, or candidate-locus datasets We can use prior evidence to define focused long-read follow-up and candidate prioritization.

Why Choose CD Genomics for Long-Read Biomarker Discovery

PacBio and ONT under one solution framework. Our LongSeq positioning is platform-complementary. We select PacBio HiFi or ONT according to the candidate structure, required molecular span, native-modification needs, and downstream analysis rather than treating one platform as universally superior.

End-to-end wet-lab and bioinformatics support. Projects can include library preparation, sequencing, custom analysis, and biological interpretation. This is important for biomarker discovery because the decisive step is often the connection between the molecular feature and the original research hypothesis, not simply the generation of raw reads.

Multiple biomarker evidence layers. Our long-read capabilities cover whole-genome variation, structural variants, haplotype phasing, full-length transcripts, single-cell isoforms, and native modification-aware workflows. A discovery program can therefore follow the candidate across genomic and transcriptomic layers without being confined to one assay type.

Clear interpretation boundaries. We distinguish candidate discovery from validation and regulated clinical qualification. Where short-read sequencing, targeted validation, or another method is more appropriate for a given stage, we can design the long-read component as a focused evidence-generating step rather than overspecifying long-read sequencing for every sample.

Independent Published Example: Long-Read Genetic Barcodes Track AML Cell Populations

Background. Penter and colleagues developed a long-read sequencing workflow to recover genetic and transcriptomic features from single-cell cDNA libraries, including somatic mutations, mitochondrial mutations, fusion genes, isoforms, T-cell receptors, and chimeric antigen sequences. The study demonstrates how long reads can add genotype-level information to cell-state profiles when short-read single-cell data alone do not provide sufficient molecular resolution.

Methods. The authors applied targeted long-read sequencing to AML/MDS single-cell cDNA libraries. Across nine patients, they targeted 11 recurrently mutated AML/MDS-associated genes and analyzed 18,097 genotyped profiles. They also used mitochondrial variants as natural genetic barcodes to distinguish and longitudinally track cell populations.

published long-read AML mitochondrial mutation tracking example from Penter et al Figure 5Figure 5 from Penter et al., Nature Communications 2024, CC BY 4.0. Long-read detection of mitochondrial genetic barcodes for donor-recipient and leukemic cell tracking.

Results. In Figure 5, the authors showed that mitochondrial mutations detected from single-cell RNA-seq libraries with long-read sequencing could distinguish donor- and recipient-derived cells and support leukemic tracking. The 10685G>A mitochondrial mutation remained detectable in myeloid progenitor cells during treatment even as their frequency decreased; the authors described it as a potential disease marker for that AML case. In another comparison, 1,855 recipient-derived and 532 donor-derived cells were annotated concordantly using mitochondrial markers and expressed-SNP analysis, with disagreement in only 14 cells. These observations illustrate how a long-read-resolved molecular feature can move from sequence detection to cell-population tracking and candidate-marker interpretation.

Conclusion. This is an independent published example, not a CD Genomics customer project or a performance guarantee. It shows the biomarker-discovery principle behind this solution: long reads can resolve molecular features that connect genotype, transcript structure, clonal identity, and cell state, creating more specific candidates for downstream research validation.

Source: Penter L, Borji M, Nagler A, et al. Integrative genotyping of cancer and immune phenotypes by long-read sequencing. Nature Communications. 2024;15:32. Figure 5, CC BY 4.0.

FAQs

Sample Deliverables

A typical biomarker-discovery report can include:

  1. A candidate evidence matrix ranking features by molecular class, sample recurrence, effect direction, and annotation context.
  2. Structural-variant locus views with breakpoint-spanning long reads and phased neighboring variants.
  3. Full-length isoform or fusion-transcript structures with condition- or cell-type-specific support.
  4. Methylation or modification tracks aligned to the candidate locus when native-molecule data are part of the study.
  5. A concise candidate shortlist with evidence summary, interpretation boundaries, and recommended validation route.

sample long-read biomarker discovery report with candidate evidence matrix phased locus isoform and methylation views

References

  1. Penter L, Borji M, Nagler A, et al. Integrative genotyping of cancer and immune phenotypes by long-read sequencing. Nature Communications. 2024;15:32.
  2. Inamo J, Suzuki A, Ueda MT, et al. Long-read sequencing for 29 immune cell subsets reveals disease-linked isoforms. Nature Communications. 2024;15:4285.
  3. Liu YH, Luo C, Golding SG, Ioffe JB, et al. Tradeoffs in alignment and assembly-based methods for structural variant detection with long-read sequencing data. Nature Communications. 2024;15:2447.
  4. Ahsan MU, Gouru A, Chan J, Zhou W, et al. A signal processing and deep learning framework for methylation detection using Oxford Nanopore sequencing. Nature Communications. 2024;15:1448.

For Research Use Only. Not for use in diagnostic or clinical procedures.

Get Your Instant Quote