
Biomarker discovery often stalls after an initial association because the underlying molecular feature is more complex than a short-read signal suggests. Long-read sequencing expands the searchable space by resolving complete structural variants, phased haplotypes, full-length transcript isoforms, fusion transcripts, and native modification patterns that can be obscured by fragmented sequencing.
CD Genomics provides a biomarker discovery solution built around complementary PacBio HiFi and Oxford Nanopore Technologies (ONT) workflows. We help research, biotech, and biopharma teams move from an unresolved genomic or transcriptomic signal to a prioritized, validation-ready set of molecular candidates with sequence, structure, phase, expression, and modification evidence matched to the research question.
Solution highlights
Many discovery programs begin with short-read WGS, RNA-seq, single-cell RNA-seq, genotyping arrays, or targeted panels. These technologies are efficient for broad screening, but some candidate regions remain unresolved because the biological feature is longer than the read, embedded in repeats, structurally rearranged, or distributed across multiple transcript exons. In those settings, deeper short-read sequencing may add coverage without adding the missing molecular context.
Long reads change the unit of evidence. A single molecule can span a structural-variant breakpoint, carry multiple heterozygous variants for phasing, cover an entire transcript isoform, or retain native modification information. This allows the discovery team to ask not only whether a locus is associated with a phenotype, but also which molecular configuration is present and whether that configuration is linked to a specific transcript, haplotype, or regulatory state.
We do not position long-read sequencing as a universal replacement for short-read discovery. Large cohorts focused on common SNVs or gene-level expression may still be more efficiently screened with high-throughput short-read methods. Long-read sequencing is most valuable when molecular structure, phase, isoform identity, repetitive context, or modification state changes the interpretation of the candidate signal.
Candidate biomarkers can be driven by insertions, deletions, inversions, duplications, translocations, repeat expansions, or compound rearrangements that are difficult to reconstruct from short fragments. Long reads can span breakpoints and repetitive flanking sequence, providing direct evidence for the architecture of the event. For projects centered on genome-wide discovery, our Human Whole Genome Sequencing workflow provides long-read coverage for genome-scale variant analysis. When the central question is a complex rearrangement, Human Genome Structural Variation Detection provides a focused route for SV discovery and interpretation.
A candidate may depend on whether variants occur in cis or trans, whether a regulatory change sits on the same molecule as a coding variant, or whether multiple variants define a risk- or response-associated haplotype. Long reads phase variants directly across longer genomic intervals, allowing candidate interpretation at the haplotype level rather than as isolated variant calls. This is particularly useful for highly polymorphic loci, duplicated regions, pharmacogenomic genes, and studies where allele-specific effects are biologically relevant.
Gene-level abundance can hide changes in transcript structure. A biologically important signal may be an isoform switch, novel splice junction, alternative transcription start or end site, intron-retaining transcript, or fusion transcript rather than a change in total gene expression. Our Full-Length Transcript Sequencing (Iso-Seq) service captures complete transcript structures for isoform discovery and comparison. Long-read transcriptomics can therefore turn an ambiguous gene-level association into a defined transcript-level candidate that can be carried into targeted validation.
Single-cell expression data are powerful for defining cell states, but short-read libraries can lose the full molecular structure needed to connect cell identity with isoforms, fusion genes, or other transcript-linked genetic features. Our Single-Cell Full-Length Transcriptome Sequencing workflow adds full-length transcript information to cell-resolved studies, supporting candidate discovery in heterogeneous systems where a signal may be restricted to a rare cell population or specific cellular state.
Long-read single-molecule sequencing can connect sequence context and epigenetic state on the same molecules. For DNA-focused studies, Long-Read Sequencing of DNA Methylation supports methylation-aware analysis alongside long-range sequence information. For RNA-focused projects where native transcript information and RNA modification patterns are central to the hypothesis, Nanopore Direct RNA Sequencing can preserve molecular features that are not retained in conventional cDNA-based workflows.

The best long-read platform depends on the candidate biomarker class and the evidence required to prioritize it. We select the platform around the research question rather than forcing every project into the same workflow.
| Research need | PacBio HiFi | Oxford Nanopore | Typical decision logic |
| High-confidence genome variant discovery | Strong fit | Strong fit | Choose according to locus complexity, required read span, and project design. |
| Complex SV and repeat-spanning analysis | Strong fit | Strong fit, especially when very long molecules add value | Use the platform that best resolves the physical span of the event. |
| Haplotype-resolved analysis | Strong fit | Strong fit | Long molecules allow linked variants to be interpreted together. |
| Full-length transcript isoforms | Strong fit through Iso-Seq | Strong fit through cDNA or direct RNA workflows | Choose according to accuracy, native-RNA requirements, and isoform question. |
| Native DNA/RNA modification analysis | Project dependent | Particularly useful for native signal analysis | Prefer native-molecule workflows when modification state is part of the candidate definition. |
| Rapid focused follow-up | Project dependent | Flexible for targeted and adaptive strategies | Match throughput and enrichment strategy to the candidate set. |
For some biomarker programs, the most efficient design is sequential: broad discovery identifies a difficult candidate class, long-read sequencing resolves the molecular structure, and a separate orthogonal assay validates the prioritized candidate in a larger cohort. For others, long-read WGS or full-length transcriptomics is the primary discovery layer from the beginning.
The workflow image is intentionally presented as one horizontal decision path. Detailed project-specific steps remain in the accompanying text so the process can be adapted without implying a rigid one-size-fits-all protocol.

A long-read experiment can generate thousands of structural, transcript, or modification features. Biomarker discovery therefore requires more than variant calling. Our analysis focuses on converting the raw feature space into candidates that can be evaluated biologically and technically.
| Analysis layer | Examples of outputs | Biomarker-discovery value |
| Read and library QC | Read-length distribution, mapping summaries, coverage profiles | Confirms whether the dataset supports the intended feature class. |
| Genome variation | SNVs/indels where relevant, SVs, CNVs, breakpoint structures | Identifies sequence and structural candidates. |
| Haplotype analysis | Phased variants, allele-specific structures, haplotype blocks | Links variants into biologically interpretable molecular configurations. |
| Transcript structure | Full-length isoforms, splice junctions, fusion transcripts, TSS/TES patterns | Identifies transcript-level candidates hidden by gene-level counts. |
| Differential analysis | Condition-associated isoform, variant, or modification patterns | Prioritizes features associated with the experimental contrast. |
| Methylation/modification | Site- or region-level modification profiles, haplotype-linked patterns | Adds regulatory evidence to sequence-based candidates. |
| Single-cell integration | Cell-type-specific isoforms, mutations, fusions, immune-receptor features | Connects molecular candidates to cellular context. |
| Functional interpretation | Gene annotation, pathway context, domain impact, known-locus overlap | Helps rank candidates for validation rather than reporting an unfiltered list. |
Candidate prioritization can incorporate recurrence across biological replicates, effect direction, molecular plausibility, orthogonal evidence, cell-type specificity, locus complexity, and compatibility with a downstream validation assay. We distinguish exploratory associations from stronger mechanistic evidence and do not present discovery-stage candidates as validated diagnostic markers.
Long-read sequencing can resolve molecular features that explain why a disease-associated locus or expression signal behaves differently across samples. Examples include a structural rearrangement that changes gene dosage, a phased haplotype linked to expression, a disease-linked transcript isoform, or a methylation state associated with a regulatory region. These features can provide a more precise research hypothesis for downstream experiments.
Treatment-response studies often produce broad expression signatures. Long-read follow-up can ask whether the response is associated with a particular isoform, fusion, allele-specific transcript, structural variant, or epigenetic configuration. Candidate markers can then be prioritized for orthogonal testing in additional samples or experimental models.
A novel transcript isoform, fusion, or complex genomic rearrangement can alter protein domains, regulatory architecture, or pathway membership in ways that are missed by gene-level analysis. Long-read sequencing provides the full molecular sequence needed to predict open reading frames, annotate altered domains, and define candidate mechanisms for functional follow-up.
Genes with segmental duplications, pseudogenes, repeats, or extensive haplotype diversity can be difficult to characterize with short reads. Long reads can improve molecular resolution across the locus and support phased candidate analysis. Research conclusions remain project-specific and should be validated with an appropriate independent method before broader use.
In heterogeneous tissues, a candidate feature may be present only in a specific malignant, stromal, or immune cell population. Long-read single-cell workflows can connect full-length transcripts and genetic features to cell-state information, supporting research into lineage, clonal structure, immune receptor diversity, and cell-type-specific isoform usage.
The most useful output of a biomarker discovery project is not the longest feature list; it is a smaller set of candidates with enough evidence to justify the next experiment. We therefore structure projects around evidence progression.
Discovery evidence may include a recurrent SV, condition-associated isoform, phased allele, or methylation pattern. Context evidence asks whether the feature is expressed in the relevant cell type, affects a plausible gene or pathway, or co-occurs with other molecular changes. Technical evidence evaluates read support, mapping ambiguity, breakpoint or isoform structure, and sample-level reproducibility. Validation planning identifies a practical orthogonal method such as targeted sequencing, PCR-based confirmation, digital PCR, targeted expression analysis, or another assay selected for the candidate class.
This separation is important in biopharma research. A statistically interesting long-read signal is not automatically a robust biomarker. The discovery page is designed to help teams create candidates that are structurally well defined and technically traceable, while leaving clinical qualification, diagnostic cutoff definition, and regulated validation outside the scope of this research-use service.
Biomarker discovery projects can enter at different stages. Final requirements depend on the selected long-read assay, so we confirm acceptance criteria during project design rather than applying one universal input specification.
| Project entry | Typical material | Planning note |
| Genome-wide structural or haplotype discovery | High-molecular-weight genomic DNA | For current human long-read WGS workflows, ≥10 μg DNA and A260/280 of 1.8–2.0 are listed as service guidance; confirm final requirements before shipment. |
| Full-length transcript biomarker discovery | High-integrity total RNA or poly(A)+ RNA | Input and integrity targets depend on Iso-Seq or ONT RNA workflow and study design. |
| Single-cell follow-up | Viable cells, prepared single-cell material, or compatible amplified cDNA depending on workflow | Review library chemistry and barcode structure before project initiation. |
| Native methylation discovery | High-quality genomic DNA suitable for native long-read sequencing | Avoid amplification when native modification information is required. |
| Existing-data entry | PacBio, ONT, short-read, single-cell, association, or candidate-locus datasets | We can use prior evidence to define focused long-read follow-up and candidate prioritization. |
PacBio and ONT under one solution framework. Our LongSeq positioning is platform-complementary. We select PacBio HiFi or ONT according to the candidate structure, required molecular span, native-modification needs, and downstream analysis rather than treating one platform as universally superior.
End-to-end wet-lab and bioinformatics support. Projects can include library preparation, sequencing, custom analysis, and biological interpretation. This is important for biomarker discovery because the decisive step is often the connection between the molecular feature and the original research hypothesis, not simply the generation of raw reads.
Multiple biomarker evidence layers. Our long-read capabilities cover whole-genome variation, structural variants, haplotype phasing, full-length transcripts, single-cell isoforms, and native modification-aware workflows. A discovery program can therefore follow the candidate across genomic and transcriptomic layers without being confined to one assay type.
Clear interpretation boundaries. We distinguish candidate discovery from validation and regulated clinical qualification. Where short-read sequencing, targeted validation, or another method is more appropriate for a given stage, we can design the long-read component as a focused evidence-generating step rather than overspecifying long-read sequencing for every sample.
Background. Penter and colleagues developed a long-read sequencing workflow to recover genetic and transcriptomic features from single-cell cDNA libraries, including somatic mutations, mitochondrial mutations, fusion genes, isoforms, T-cell receptors, and chimeric antigen sequences. The study demonstrates how long reads can add genotype-level information to cell-state profiles when short-read single-cell data alone do not provide sufficient molecular resolution.
Methods. The authors applied targeted long-read sequencing to AML/MDS single-cell cDNA libraries. Across nine patients, they targeted 11 recurrently mutated AML/MDS-associated genes and analyzed 18,097 genotyped profiles. They also used mitochondrial variants as natural genetic barcodes to distinguish and longitudinally track cell populations.
Figure 5 from Penter et al., Nature Communications 2024, CC BY 4.0. Long-read detection of mitochondrial genetic barcodes for donor-recipient and leukemic cell tracking.
Results. In Figure 5, the authors showed that mitochondrial mutations detected from single-cell RNA-seq libraries with long-read sequencing could distinguish donor- and recipient-derived cells and support leukemic tracking. The 10685G>A mitochondrial mutation remained detectable in myeloid progenitor cells during treatment even as their frequency decreased; the authors described it as a potential disease marker for that AML case. In another comparison, 1,855 recipient-derived and 532 donor-derived cells were annotated concordantly using mitochondrial markers and expressed-SNP analysis, with disagreement in only 14 cells. These observations illustrate how a long-read-resolved molecular feature can move from sequence detection to cell-population tracking and candidate-marker interpretation.
Conclusion. This is an independent published example, not a CD Genomics customer project or a performance guarantee. It shows the biomarker-discovery principle behind this solution: long reads can resolve molecular features that connect genotype, transcript structure, clonal identity, and cell state, creating more specific candidates for downstream research validation.
Source: Penter L, Borji M, Nagler A, et al. Integrative genotyping of cancer and immune phenotypes by long-read sequencing. Nature Communications. 2024;15:32. Figure 5, CC BY 4.0.
Add long reads when the unresolved part of the signal depends on molecular structure or linkage: a complex SV, repeat-associated locus, phased haplotype, full-length isoform, fusion transcript, native methylation pattern, or cell-specific transcript structure. If the question is limited to common SNVs or gene-level expression across a very large cohort, short-read sequencing may remain the more efficient primary platform.
It can discover candidate molecular features that were not visible or could not be fully reconstructed with shorter reads. These may include previously unresolved SVs, novel full-length isoforms, fusion transcripts, phased variant combinations, or modification patterns. They remain research candidates until validated statistically, biologically, and technically.
The choice depends on the candidate class. PacBio HiFi is well suited to high-confidence sequence and isoform resolution. ONT is highly flexible for long-range molecules and native DNA/RNA signal analysis. Some projects benefit from a staged or combined strategy. We select the platform after reviewing the biological question, sample type, locus complexity, and downstream validation plan.
Yes. Existing datasets can define the long-read follow-up question. Examples include unresolved SV calls, candidate loci from GWAS/WGS, gene-level RNA-seq hits that require isoform resolution, or single-cell clusters where genotype or fusion status needs to be linked to cellular identity.
Native long-read DNA sequencing can support sequence and modification analysis on long molecules, enabling haplotype-aware or locus-aware interpretation when the study design and data quality support it. The exact analysis plan depends on platform, coverage, sample quality, and the candidate region.
No. The same discovery logic applies to immunology, rare-disease research, pharmacogenomics, neurological research, inherited traits, microbial and host-response studies, and other research areas where molecular structure or phase matters. Assay selection is adapted to the biological system and question.
We can design the discovery output to support downstream targeted confirmation and can discuss sequencing-based follow-up options. The appropriate validation method depends on the candidate type. Clinical qualification, diagnostic performance claims, and regulated clinical validation are outside the scope of this RUO solution.
A typical biomarker-discovery report can include:

References
For Research Use Only. Not for use in diagnostic or clinical procedures.