Full-length splice variant analysis, fusion gene detection, and isoform-level expression profiling by Oxford Nanopore cDNA sequencing on PromethION — PCR-cDNA and direct cDNA workflows supported
CD Genomics provides Nanopore Full-Length cDNA Sequencing on the PromethION platform, delivering 50–100+ million reads per flow cell that capture complete transcript isoforms as single contiguous reads. Using both PCR-cDNA and direct cDNA workflows, our service enables full-length isoform discovery, splice variant detection, fusion gene identification, and isoform-level quantification across any eukaryotic species — without the computational reconstruction required by short-read RNA-seq methods.
The transcriptome is not defined by gene expression levels alone — it is defined by which isoforms are expressed, in what proportions, and how splice variation, alternative promoter usage, and alternative polyadenylation reshape the functional landscape of the cell. Short-read RNA sequencing, which fragments transcripts into 50–300 bp pieces before sequencing, observes the transcriptome through fragmented windows and relies on computational inference to reconstruct full-length isoform architectures from partial evidence. As transcript complexity increases — with multi-exon skipping, mutually exclusive exons, intron retention events, and complex alternative splicing patterns — the inference problem becomes increasingly ill-posed, and short-read data systematically underestimates isoform diversity.
CD Genomics offers Nanopore Full-Length cDNA Sequencing — a high-throughput long-read RNA sequencing service that captures complete transcript isoforms as single contiguous reads. Using Oxford Nanopore Technologies (ONT) PromethION and GridION platforms with both PCR-cDNA and direct cDNA library preparation workflows, our service delivers full-length transcript sequences spanning from the 5′ cap region to the 3′ poly(A) tail, enabling direct observation of splice isoform architecture, isoform-level quantification across experimental conditions, fusion transcript discovery, and alternative splicing analysis at single-molecule resolution without computational reconstruction.
At a glance:
The human transcriptome is estimated to produce over 200,000 distinct transcript isoforms from approximately 20,000 protein-coding genes, with the majority of multi-exon genes expressing multiple splice variants that differ in their exon composition, untranslated region architecture, and coding potential. These isoforms are not functionally equivalent. Alternative splicing can alter protein domain composition, change subcellular localization signals, introduce or remove regulatory elements in untranslated regions, and produce transcripts targeted for nonsense-mediated decay. The choice between isoforms is cell-type-specific, developmentally regulated, and frequently dysregulated in disease — yet short-read RNA-seq, the most widely used transcriptome profiling method, provides only indirect evidence for these isoform-level events.
Full-length cDNA sequencing addresses this gap by observing complete transcript isoforms directly. Unlike short-read methods that fragment RNA before sequencing and then computationally reassemble the fragments into inferred transcript models, full-length cDNA sequencing captures each transcript molecule in its entirety as a single sequencing read. The isoform structure — which exons are included, which splice sites are used, where the transcript starts and ends — is read directly from the sequencing data, not inferred from fragment coverage patterns. This distinction is critical for accurate isoform-level analysis: studies consistently show that short-read methods systematically underestimate isoform diversity, miss rare isoforms, and produce ambiguous isoform assignments for genes with complex splicing patterns.
Nanopore cDNA sequencing occupies a specific position in the long-read RNA sequencing landscape. It provides the highest throughput among full-length transcript methods (50–100+ million reads per flow cell), making it the method of choice when the biological question requires deep coverage of the transcriptome for comprehensive isoform discovery and reliable quantification of differentially expressed isoforms. Unlike our Nanopore Direct RNA Sequencing service, cDNA sequencing requires reverse transcription and (in the PCR-cDNA workflow) PCR amplification, which means RNA base modification information is not preserved. The trade-off is throughput: where direct RNA sequencing provides 5–20 million reads per flow cell with full modification information, cDNA sequencing delivers an order of magnitude more reads for projects where transcript structure and expression, not modifications, are the primary question.
Nanopore Full-Length cDNA Sequencing is an Oxford Nanopore Technology-based method that sequences complete complementary DNA (cDNA) molecules derived from cellular RNA. The method captures the full-length connectivity of each transcript isoform in a single contiguous read, enabling direct observation of splice isoform architecture, transcript boundaries, and alternative processing events without the fragmentation-and-reconstruction workflow of short-read RNA-seq.
The sequencing process involves two main library preparation workflows, each optimized for different experimental priorities:
PCR-cDNA Workflow (ONT SQK-PCS111). Total RNA or poly(A)-selected RNA is used as template for reverse transcription with an oligo(dT) primer that captures the 3′ poly(A) tail. A template-switching oligonucleotide (TSO) captures the 5′ cap structure, enabling full-length first-strand cDNA synthesis that spans from the 5′ cap to the 3′ poly(A) tail. The resulting full-length cDNA is then PCR-amplified to generate sufficient material for nanopore sequencing. This workflow provides the highest throughput (typically 50–100+ million reads per PromethION flow cell) and is the preferred choice for comprehensive transcript discovery projects where maximum sequencing depth is the primary priority. The template-switching step also enables efficient capture of the 5′ transcript end, providing accurate transcription start site information.
Direct cDNA Workflow (ONT SQK-DCS111). First-strand cDNA is synthesized using an oligo(dT) primer that incorporates the motor protein adapter sequence, eliminating the need for separate adapter ligation. The cDNA is sequenced directly without PCR amplification, reducing amplification bias and preserving the quantitative relationship between transcript abundance and read count more faithfully. This workflow provides lower throughput (typically 10–30 million reads per flow cell) than PCR-cDNA but is preferred when strand-of-origin information is required (the direct cDNA method is strand-aware, while PCR-cDNA with template-switching can lose strand information), when minimizing amplification bias is critical for quantitative accuracy, or when limited PCR cycles are desired to preserve the representation of GC-rich transcripts.
In both workflows, the prepared library is loaded onto a PromethION or GridION flow cell containing an array of protein nanopores embedded in an electrically resistant membrane. A voltage is applied across the membrane, driving negatively charged cDNA molecules through the nanopores. As each cDNA molecule translocates through a pore, it modulates the ionic current in a sequence-dependent manner. The current is measured at kHz frequency and basecalled in real time using ONT Dorado, producing full-length cDNA reads that span the complete transcript isoform. The key feature shared by both workflows is that no fragmentation step is involved — every read represents an intact full-length transcript molecule, providing direct evidence for the connectivity between exons, the selection of splice sites, and the boundaries of the transcript.
Splice isoform structure is observed directly in each sequencing read. Exon connectivity, splice junction usage, and alternative exon inclusion levels are determined from alignment of full-length reads to the reference genome, without computational isoform inference or assembly. This eliminates the isoform ambiguity that arises from short-read fragment-based methods, where reads spanning individual splice junctions must be assembled into complete transcript models — a process that becomes increasingly unreliable as transcript complexity increases. Complex alternative splicing events — multi-exon skipping, mutually exclusive exons, intron retention, and alternative 5′/3′ splice site selection — are resolved unambiguously from full-length read alignments.
With 50–100+ million reads per PromethION flow cell (PCR-cDNA workflow), the sequencing depth is sufficient to detect low-abundance transcripts, rare splicing events, and isoforms expressed at a few copies per cell that are missed by lower-throughput long-read methods. The comprehensive coverage also supports robust detection of novel transcripts not present in reference annotations, including previously unannotated isoforms of known genes, intergenic transcripts, and fusion transcripts arising from genomic rearrangements or read-through transcription.
Isoform-level expression quantification from full-length cDNA reads avoids the multi-mapping ambiguity that limits short-read isoform quantification. Each read is assigned to a specific full-length isoform based on its complete exon composition, eliminating the need for statistical deconvolution of ambiguous read assignments. The resulting isoform expression estimates are more accurate than short-read-based methods, particularly for isoforms that share extensive exon structure, and enable confident identification of condition-specific isoform switching events.
We offer both PCR-cDNA and direct cDNA library preparation workflows, with project-specific recommendations based on sample availability, throughput requirements, strand-awareness needs, and bias sensitivity. For projects requiring maximum transcript discovery depth, the PCR-cDNA workflow delivers 50–100+ million reads per flow cell. For projects prioritizing quantitative accuracy and strand information, the direct cDNA workflow provides amplification-free sequencing with lower throughput but reduced bias. Our project scientists provide method selection guidance to match the library preparation approach to each project's specific biological question.
We have established and optimized protocols for sequencing 10x Genomics single-cell cDNA libraries on the ONT platform, enabling single-cell full-length transcript isoform discovery at scale. Our low-input protocols support bulk RNA inputs as low as 100 ng total RNA, and we provide feasibility assessments for challenging sample types including FFPE-derived RNA, clinical biopsies, and immunoprecipitated RNA from RNA-binding protein studies. For single-cell projects, our Single-Cell Full-Length Transcriptome Sequencing service provides end-to-end support from 10x library preparation through ONT sequencing and isoform-level analysis.
We provide comprehensive project support including RNA extraction and QC guidance, library preparation with the optimal ONT kit for each project, PromethION or GridION sequencing with real-time monitoring, dedicated isoform-level bioinformatics analysis, and custom visualization for manuscript preparation. Our project scientists offer integrated analysis guidance for researchers combining cDNA sequencing data with other approaches in their multi-platform RNA analysis strategy.
Our service is designed for integration with orthogonal RNA analysis methods. cDNA sequencing data can be combined with Nanopore Direct RNA Sequencing from the same RNA samples to link isoform structure with RNA modification status, or with PacBio RNA Sequencing for HiFi-accuracy validation of specific isoforms of interest. Our bioinformatics pipelines support cross-platform data integration, enabling multi-faceted transcriptome characterization from a single project.
The table below provides a direct comparison of Nanopore Full-Length cDNA Sequencing against the three other major RNA-seq approaches, enabling researchers to evaluate which method best matches their experimental requirements. Nanopore cDNA sequencing is shown for both the PCR-cDNA and direct cDNA workflows, as the choice between them significantly affects throughput, bias profile, and information content.
| Feature | Nanopore PCR-cDNA Sequencing | Nanopore Direct cDNA Sequencing | PacBio Iso-Seq (Kinnex) | Short-Read RNA-seq (Illumina) |
| Sequencing template | PCR-amplified full-length cDNA | Full-length cDNA (no amplification) | PCR-amplified full-length cDNA (HiFi consensus) | PCR-amplified cDNA fragments |
| Throughput per run | Very high (50–100+ million reads) | Moderate (10–30 million reads) | Moderate (2–60 million reads per SMRT Cell with Kinnex) | Very high (100–500+ million reads per lane) |
| Read length | 500–5,000+ nt (full-length cDNA; limited by RT processivity and PCR) | 500–3,000+ nt (full-length cDNA; limited by RT processivity) | 1,000–10,000+ nt (PCR-amplified, size-selected) | 50–300 bp (fragmented) |
| Per-base accuracy | Moderate (~Q10–Q15 raw; consensus improves accuracy) | Moderate (~Q10–Q15 raw) | High (Q30+ circular consensus) | Very high (Q30+) |
| RT/PCR bias | RT + PCR bias (amplification skew, GC bias, template-switching artifacts) | RT bias only (premature termination at structured regions, GC-dependent errors) | RT + PCR bias (similar to ONT PCR-cDNA) | RT + PCR + fragmentation bias (ligation, PCR duplication, GC-dependent amplification) |
| Native base modifications preserved | ✘ Lost during RT-PCR | ✘ Lost during RT | ✘ Lost during RT-PCR | ✘ Lost during RT-PCR + fragmentation |
| 5′ to 3′ transcript coverage | Full-length reads (5′ cap captured by TSO; some 5′ bias from RT drop-off) | Full-length reads (3′-biased; RT initiates at poly(A) tail, incomplete 5′ coverage) | Full-length reads (5′ and 3′ bias from template-switching) | 3′-biased (poly(A) selection + fragmentation) |
| Single-molecule information | Per-read full-length transcript sequence (single-molecule, but PCR duplicates present) | Per-read full-length transcript sequence (single-molecule, no PCR duplicates) | Consensus sequence from multiple passes (not single-molecule for accuracy) | Not single-molecule (ensemble averages from fragment coverage) |
| Single-cell compatibility | ✓ Established protocols for 10x Genomics cDNA and Smart-seq2 | ✓ Compatible with 10x single-cell cDNA | ✓ MAS-seq for 10x cDNA (established protocols) | ✓ Standard scRNA-seq platforms (10x, Drop-seq, Smart-seq2) |
| Strand information | ~ Limited (TSO method may lose strand specificity) | ✓ Strand-aware (direct cDNA preserves strand orientation) | ~ Limited (depends on library prep) | ✓ Standard strand-specific RNA-seq protocols available |
| RNA input requirement | 100 ng–1 μg total RNA (PCR amplification enables lower input) | 100 ng–1 μg total RNA (higher input required than PCR-cDNA) | 100 ng–1 μg total RNA | 10–500 ng total RNA |
| Bioinformatics maturity | Mature (bambu, Isosceles, Flair, isoform quantification tool; extensive community tooling) | Mature (same tools as PCR-cDNA; strand-aware analysis adds value) | Mature (IsoSeq3, Cupcake, SQANTI) | Very mature (STAR, RSEM, Salmon, Kallisto, DEXSeq) |
| Best suited for | High-depth isoform discovery, rare transcript detection, transcriptome-wide isoform quantification and differential analysis, fusion gene discovery, comprehensive transcript annotation | Strand-aware isoform analysis, amplification-bias-sensitive quantification, integrative projects combining cDNA with direct RNA data | High-accuracy isoform validation, de novo transcript annotation requiring consensus accuracy, fusion transcript validation, targeted isoform sequencing | High-throughput gene-level expression quantification, differential expression analysis, small RNA analysis, large-scale population transcriptomics |
The PCR-cDNA workflow's deep sequencing depth (50–100+ million reads per flow cell) enables detection of low-abundance isoforms, while the direct cDNA workflow's strand-aware protocol provides accurate annotation of antisense transcripts and non-coding RNA isoforms — a capability not available from most other long-read RNA sequencing services.
Our Nanopore Full-Length cDNA Sequencing service is part of a comprehensive transcriptome analysis ecosystem. The following complementary service modules support projects across different research objectives and provide integrated multi-platform solutions.
Full-Length Transcriptome Profiling provides an end-to-end solution for transcript discovery, isoform quantification, and functional annotation across any eukaryotic species. This module integrates Nanopore cDNA sequencing with comprehensive bioinformatics analysis optimized for both model and non-model organisms, and can be combined with downstream validation using PacBio RNA Sequencing for HiFi-accuracy confirmation of specific isoforms.
Nanopore Direct RNA Sequencing sequences native RNA molecules directly without reverse transcription or PCR amplification, preserving every base modification (m6A, m5C, Ψ, inosine) and enabling poly(A) tail length measurement. We recommend dRNA-seq as the complementary method to cDNA sequencing for projects requiring both deep isoform coverage and modification information — the same RNA sample can be split for parallel cDNA and direct RNA sequencing, providing orthogonal data from the same biological material.
Single-Cell Full-Length Transcriptome Sequencing provides specialized support for single-cell isoform discovery projects, including 10x Genomics single-cell cDNA library preparation, ONT sequencing, and single-cell-optimized bioinformatics analysis. This module handles the specific challenges of single-cell long-read data, including molecular barcode/barcode processing, cell-barcode-aware isoform quantification, and integration with single-cell gene expression data.
Oxford Nanopore Sequencing Data Analysis provides dedicated bioinformatics support for full-length cDNA sequencing data, including basecalling with Dorado super-accuracy models, full-length isoform detection and quantification (bambu, Isosceles, Flair), differential transcript usage analysis, splice junction quantification, fusion transcript discovery, and integrated multi-platform report generation combining cDNA sequencing data with data from other long-read or short-read methods.
Total RNA is extracted from the sample using methods optimized for RNA integrity and purity. RNA quantity and quality are assessed by microfluidic electrophoresis (RIN ≥ 8 recommended for standard protocols). Poly(A) selection using oligo-dT beads is performed for mRNA-focused studies; rRNA depletion is available for total RNA analysis including non-coding transcripts. DNase treatment eliminates genomic DNA contamination. For single-cell cDNA sequencing projects, cDNA is generated from 10x Genomics single-cell libraries using the 10x pipeline and then converted to ONT-compatible libraries.
For the PCR-cDNA workflow (ONT SQK-PCS111), reverse transcription is performed using an oligo(dT) primer that captures the 3′ poly(A) tail. The template-switching oligonucleotide (TSO) captures the 5′ cap structure of full-length mRNAs, enabling the reverse transcriptase to complete full-length first-strand cDNA synthesis. The resulting cDNA contains the complete transcript sequence from the 5′ cap region through the coding sequence to the 3′ poly(A) tail, with adapters at both ends for subsequent PCR amplification and sequencing. For the direct cDNA workflow (ONT SQK-DCS111), the oligo(dT) primer includes the motor protein adapter sequence, enabling direct sequencing of the first-strand cDNA without PCR amplification. Reverse transcription is performed with a thermostable RT enzyme optimized for full-length cDNA synthesis, minimizing premature termination at structured RNA regions.
Figure 1. Nanopore Full-Length cDNA Sequencing workflow: from RNA extraction and quality assessment through full-length cDNA synthesis (PCR-cDNA or direct cDNA library preparation) and Oxford Nanopore sequencing on PromethION or GridION platforms.
For PCR-cDNA: The full-length cDNA is PCR-amplified using primers specific to the adapters introduced during reverse transcription, with the number of PCR cycles optimized to balance yield against amplification bias (typically 12–18 cycles). Sequencing adapters containing the motor protein and tether are ligated to the amplified cDNA. For direct cDNA: The first-strand cDNA, already containing the motor protein adapter incorporated during reverse transcription, is directly prepared for sequencing without amplification. In both workflows, the library is purified to remove short fragments, unligated adapters, and other contaminants before loading onto the flow cell.
The prepared library is loaded onto a PromethION or GridION flow cell. The motor protein bound to the cDNA adapter processively unwinds and translocates the cDNA molecule through the nanopore at approximately 400–450 bases per second. As each nucleotide passes through the pore, it modulates the ionic current in a sequence-dependent manner. The current is measured at kHz frequency, generating a continuous signal trace that encodes the complete nucleotide sequence of the cDNA molecule. Basecalling is performed in real time using ONT Dorado with super-accuracy (SUP) models, producing high-quality FASTQ reads for downstream analysis. The sequencing run can be monitored in real time, with run termination decisions based on accumulated data yield, target coverage, and project-specific quality metrics.
Basecalled reads are filtered by quality score and read length, then aligned to the reference genome or transcriptome using splice-aware aligners (minimap2 with splice parameters, or uLTRA for more sensitive alignment of full-length transcripts). Downstream analysis includes: isoform detection and quantification using dedicated long-read tools (bambu for reference-based isoform discovery and quantification, Isosceles for single-cell-resolution analysis, Flair for full-length isoform assembly and quantification, or isoform quantification tool for comprehensive isoform analysis with splice junction support), differential transcript usage (DTU) analysis (DRIMSeq or satuRn), splice junction quantification and alternative splicing event analysis (SUPPA2 or MAJIQ), fusion transcript detection (JAFFAL or LongGF), and functional annotation of novel isoforms including coding potential assessment, conserved domain identification, and comparison with reference annotations.
| Analysis Feature | Standard Package | Advanced Package |
| Dorado super-accuracy (SUP) basecalling with RNA-specific models | ✓ | ✓ |
| Read QC, filtering (minimum read length, quality score, adapter trimming) | ✓ | ✓ |
| Splice-aware alignment to reference genome/transcriptome (minimap2) | ✓ | ✓ |
| Full-length transcript isoform detection and quantification (bambu) | ✓ | ✓ |
| Gene-level and isoform-level expression quantification and normalization | ✓ | ✓ |
| Isoform novelty classification (known, novel in catalog, novel not in catalog) | ✓ | ✓ |
| Differential transcript usage (DTU) analysis between conditions | — | ✓ |
| Alternative splicing event quantification and classification (SUPPA2) | — | ✓ |
| Fusion transcript detection and characterization (JAFFAL, LongGF) | — | ✓ |
| Single-cell isoform analysis with cell-barcode-aware quantification (Isosceles) | — | ✓ |
| Functional annotation of novel isoforms (coding potential, domain analysis) | — | ✓ |
| Multi-platform data integration report (cDNA + DRS + PacBio RNA comparison) | — | ✓ |
| Category | PCR-cDNA Workflow (Recommended) | Direct cDNA Workflow |
| Sample type | Total RNA (eukaryotic), poly(A)+ RNA, single-cell cDNA libraries (10x Genomics), FFPE RNA; mammalian, plant, fungal samples | Total RNA (eukaryotic), poly(A)+ RNA, single-cell cDNA libraries; mammalian, plant, fungal samples |
| Minimum input | 100 ng–1 μg total RNA (PCR amplification enables lower input); 10–50 ng μg poly(A)+ RNA | 250 ng–1 μg total RNA (higher input required without PCR amplification) |
| RNA integrity | RIN ≥ 8 recommended (RIN ≥ 7 minimum); degraded RNA assessed case-by-case | RIN ≥ 8 recommended (higher integrity required for optimal yield without PCR) |
| Recommended depth | 50–100+ million reads per sample (PromethION flow cell) | 10–30 million reads per sample (PromethION flow cell) |
| Single-cell input | Amplified 10x cDNA library (≥1 ng total, quantified by Qubit hsDNA) | Amplified or unamplified 10x cDNA (direct cDNA requires higher cDNA mass) |
| Strand information | Limited (TSO-based method may lose strand specificity) | ✓ Strand-aware (direct cDNA preserves strand orientation) |
| Shipping | Overnight on dry ice (RNA or tissue); RNA in RNase-free water or storage buffer; see sample submission guidelines | |
| QC Parameter | Minimum Requirement | Recommended Target |
| Read Q-score (after basecalling) | Q7 (median) | ≥Q10 (median) |
| Read length N50 | 500 nt | ≥1,000 nt (PCR-cDNA) / ≥800 nt (direct cDNA) |
| Aligned read proportion | 60% | ≥80% |
| Full-length read proportion (aligned 5′ to 3′) | 40% | ≥60% (PCR-cDNA with TSO) / ≥40% (direct cDNA) |
| Isoform detection sensitivity (known genes) | 60% | ≥80% at 30× coverage |
Validated dual-workflow cDNA sequencing on PromethION and GridION
We have established and optimized both PCR-cDNA (SQK-PCS111) and direct cDNA (SQK-DCS111) library preparation workflows across multiple PromethION and GridION instruments, with validated protocols for diverse eukaryotic species (human, mouse, rat, zebrafish, Arabidopsis, rice, yeast, fungi), single-cell cDNA libraries from multiple 10x chemistries, and challenging sample types including FFPE-derived RNA and biopsy material. Our sequencing yields are benchmarked against published SG-NEx consortium performance data, and we provide detailed feasibility assessments for challenging projects.
Method selection guidance matched to biological questions
We recognize that no single library preparation method is optimal for every project. Our project scientists provide evidence-based recommendations for PCR-cDNA vs direct cDNA workflows based on: whether maximum throughput or minimal bias is the primary priority, whether strand-of-origin information is required, the RNA input amount available, the target transcript expression range (rare transcripts benefit from PCR-cDNA depth), and the planned integration with other data types (e.g., parallel direct RNA sequencing). This method selection guidance ensures that each project receives the library preparation strategy best suited to its specific biological question.
Comprehensive bioinformatics from isoform discovery to biological interpretation
Our bioinformatics pipeline covers the complete analysis journey from raw basecalled reads to biological interpretation. Standard analysis includes full-length isoform detection, quantification, and novelty classification using validated tools benchmarked against the SG-NEx and Isosceles resources. Advanced analysis adds differential transcript usage, alternative splicing event quantification, fusion transcript detection, single-cell isoform analysis, and functional annotation of novel isoforms — providing publication-ready results that connect transcript structural changes with their biological context.
Multi-platform integration for comprehensive transcriptome characterization
Our Nanopore Full-Length cDNA Sequencing service is designed for integration with our complete RNA analysis portfolio. Researchers can combine cDNA sequencing with Nanopore Direct RNA Sequencing for epitranscriptome analysis, PacBio RNA Sequencing for HiFi-accuracy isoform validation, Long-Read Sequencing of RNA Methylation for targeted modification analysis, and Oxford Nanopore Sequencing Data Analysis for comprehensive cross-platform bioinformatics support — all within a single coordinated project framework.
Proven track record in full-length transcript sequencing across diverse applications
Our full-length cDNA sequencing service has supported published research across cancer transcriptomics, neurobiology, plant biology, developmental biology, and comparative genomics. Our team's expertise spans from experimental design through bioinformatics interpretation, providing comprehensive support for researchers at every stage of their transcriptome analysis project. We are committed to delivering the sequencing depth, data quality, and analytical rigor required for high-impact transcriptome research.
Kabza M, Ritter A, Byrne A, Sereti K, Le D, Stephenson W, Sterne-Weiler T. Accurate long-read transcript discovery and quantification at single-cell, pseudo-bulk and bulk resolution with Isosceles. Nature Communications. 2024;15:7316. (CC BY 4.0)
Long-read RNA sequencing technologies, including Nanopore cDNA sequencing, offer the ability to sequence full-length transcript isoforms, but the computational analysis of long-read data presents unique challenges. Technical noise, variable read quality, and the complexity of isoform-level quantification — particularly at the single-cell level — require specialized computational methods that can distinguish true biological isoform diversity from technical artifacts. Existing tools (Bambu, isoform quantification tool, ESPRESSO) each have limitations in sensitivity and accuracy, especially for single-cell data where per-cell coverage is limited. The authors developed Isosceles, a computational toolkit designed to address these challenges by using acyclic splice-graph representations of gene structure combined with an Expectation-Maximization (EM) algorithm for accurate isoform detection and quantification from Nanopore long-read RNA sequencing data.
Isosceles was developed using a splice-graph-based approach in which gene structure is represented as a directed acyclic graph, with exons as vertices and splice junctions as edges. Both known transcripts from reference annotations and de novo transcripts identified from long-read data are incorporated into the graph. Transcript compatibility counts (TCCs) are calculated by aligning each long read to the graph and determining which transcripts are compatible with the observed splice junctions. An EM algorithm is then applied to the TCCs to estimate transcript-level expression. The method was benchmarked using synthetic RNA spike-in data with known isoform composition, as well as bulk and single-cell Nanopore cDNA sequencing data from human cell lines and mouse brain. Performance was compared against Bambu, isoform quantification tool, and ESPRESSO across multiple metrics including isoform detection sensitivity, quantification accuracy, and computational efficiency.
Figure 2. Isosceles architecture and key results: splice-graph-based transcript discovery and quantification from Nanopore long-read cDNA sequencing data, with validation across single-cell, pseudo-bulk, and bulk resolutions. From Kabza M, Ritter A, Byrne A, et al. (2024, Nature Communications, CC BY 4.0).
This study provides a validated computational framework for accurate transcript discovery and quantification from Nanopore full-length cDNA sequencing data, with demonstrated performance across bulk, pseudo-bulk, and single-cell resolutions. The Isosceles toolkit addresses a critical need in the long-read transcriptomics field — the challenge of accurate isoform-level quantification from noisy long-read data — and achieves significant improvements over existing methods in both sensitivity and accuracy. The biological discoveries enabled by reanalysis of mouse brain single-cell data demonstrate that improved computational methods can extract additional biological insight from existing Nanopore cDNA sequencing datasets. These findings validate our bioinformatics pipeline design for full-length cDNA sequencing projects, which includes Isosceles as a key component of the Advanced analysis package for single-cell isoform analysis and demonstrates the value of ongoing computational method development for extracting maximal biological information from Nanopore full-length cDNA sequencing data.
CD Genomics provides free project consultation to help determine the optimal RNA sequencing strategy for your specific research question. Contact our scientists to discuss your project requirements.
Nanopore cDNA Sequencing requires reverse transcription of RNA into cDNA before sequencing, which means all native RNA base modification information (m6A, m5C, Ψ, inosine) is lost during the RT step. However, cDNA sequencing delivers 5–10 times more reads per flow cell (50–100+ million vs 5–20 million for direct RNA sequencing), making it the preferred choice for deep isoform discovery, rare transcript detection, and transcriptome-wide expression profiling. Direct RNA Sequencing (dRNA-seq) sequences native RNA directly, preserving all base modifications and enabling poly(A) tail measurement, but at lower throughput. The choice depends on whether modification information or maximum sequencing depth is the primary priority for your biological question. For projects where both depth and modification analysis are important, the same RNA sample can be split for parallel cDNA and direct RNA sequencing.
No. Because Nanopore cDNA Sequencing requires reverse transcription of RNA into cDNA, all RNA base modifications are lost during the first-strand synthesis step. The sequencing signal is generated from the cDNA molecule, which does not retain modification information. If RNA modification detection is required for your project, we recommend our Nanopore Direct RNA Sequencing service, which sequences the native RNA molecule directly and preserves all base modifications. For projects requiring both modification analysis and deep isoform coverage, we offer parallel cDNA and direct RNA sequencing from the same RNA sample.
PCR-cDNA (ONT SQK-PCS111) uses reverse transcription with a template-switching oligonucleotide to capture the 5′ cap, followed by PCR amplification. This workflow provides the highest throughput (50–100+ million reads per flow cell) and captures 5′ transcript ends efficiently, but introduces PCR amplification bias and may lose strand-of-origin information. Direct cDNA (ONT SQK-DCS111) uses an oligo(dT) primer containing the motor protein adapter sequence, and the cDNA is sequenced directly without PCR amplification. This workflow provides lower throughput (10–30 million reads per flow cell) but avoids amplification bias and preserves strand information. Our project scientists recommend the optimal workflow based on your project's throughput requirements, need for strand information, and RNA input availability.
Nanopore cDNA sequencing typically delivers 50–100+ million reads per PromethION flow cell with the PCR-cDNA workflow, and 10–30 million reads with the direct cDNA workflow. The number of reads required for comprehensive isoform analysis depends on the transcriptome complexity of the species, the number of samples being multiplexed, and whether the goal is discovery of novel isoforms or quantification of known isoforms. For a typical mammalian transcriptome project focused on isoform discovery and quantification, we recommend 30–50 million reads per sample for bulk analysis. For single-cell projects, coverage requirements are typically 500,000–2 million reads per cell, depending on the number of cells and the complexity of splice patterns being analyzed. Our project scientists will help determine the optimal sequencing depth for your specific experimental design.
Yes. We have validated protocols for sequencing 10x Genomics single-cell cDNA libraries on the ONT PromethION platform, compatible with both 3′ (v2, v3, Next GEM) and 5′ single-cell chemistries. The 10x cDNA is converted to an ONT-compatible library using established protocols (SQK-PCS111 for PCR-based conversion or dedicated single-cell long-read kits). Our bioinformatics pipeline includes cell-barcode-aware analysis using Isosceles or Bambu-Clump, which are specifically designed for the challenges of single-cell long-read data including lower per-cell coverage and higher technical noise. For comprehensive single-cell transcript discovery projects, see our Single-Cell Full-Length Transcriptome Sequencing service page for detailed protocols and project specifications.
We deploy a validated suite of long-read-specific bioinformatics tools. For standard isoform detection and quantification, we use bambu for reference-based isoform discovery and quantification with novelty classification. For single-cell projects, we use Isosceles with its cell-barcode-aware Expectation-Maximization algorithm, which achieves Spearman correlation of 0.96 against ground-truth data. For differential transcript usage analysis between conditions, we use DRIMSeq or satuRn. Fusion transcript detection is performed using JAFFAL or LongGF. All pipelines are benchmarked against published standards including the SG-NEx consortium dataset and are regularly updated as new tools become available.
Standard turnaround time for a Nanopore full-length cDNA sequencing project (including library preparation, sequencing, and basic bioinformatics analysis) is approximately 4–6 weeks from sample receipt to final data delivery, depending on the number of samples, library preparation workflow selected, and sequencing depth required. Expedited timelines can be arranged for time-sensitive projects. Projects with the Advanced Bioinformatics package typically require an additional 1–2 weeks for the extended analysis. We recommend contacting our project management team during project initiation to establish a timeline that matches your submission deadline.
Yes. Integrating Nanopore full-length cDNA sequencing with short-read RNA-seq data from the same RNA samples is a powerful hybrid approach. Short-read data provides deep coverage for precise gene-level expression quantification, while full-length cDNA reads resolve the isoform structures that short reads alone cannot distinguish. Our bioinformatics pipelines support hybrid analysis workflows, including short-read-based isoform quantification guided by long-read-derived isoform models, and integrated visualization of both data types in genome browsers. We recommend this hybrid approach for projects where both high-throughput expression quantification and comprehensive isoform resolution are required.
In addition to standard FASTQ, BAM, and expression matrix deliverables, our Advanced package provides optional multi-platform integration reports that correlate cDNA isoform data with direct RNA modification analysis or PacBio HiFi isoform validation from the same biological samples.
1. Basecalled FASTQ files from Dorado super-accuracy (SUP) basecalling with RNA-specific models, including per-read quality metrics, adapter-trimmed sequences, and read length distributions — suitable for downstream analysis and public data deposition in SRA or GEO
2. Genome-aligned BAM files with splice-aware alignments (minimap2), transcript isoform annotations, and splice junction labels — ready for visualization in IGV or UCSC Genome Browser showing full-length read coverage across gene models with isoform-specific read assignments
3. Transcript isoform annotation and quantification report including full-length isoform structures with exon-intron boundaries, isoform-level expression matrices (TPM/read counts) for all detected isoforms, isoform novelty classification (known/novice in catalog/novel not in catalog), and comparison with reference annotations
4. Isoform-level analysis report covering differential transcript usage (DTU) analysis between experimental conditions with statistical testing, alternative splicing event quantification with event-level significance, fusion transcript calls with breakpoint coordinates and supporting read evidence, and functional annotation of novel isoforms including coding potential and domain identification
5. Integrated multi-platform report (optional, for projects combining cDNA sequencing with other methods) correlating isoform expression with RNA modification status from parallel direct RNA sequencing, or providing cross-platform isoform validation comparing cDNA and PacBio RNA sequencing data for isoforms of interest
References
For research use only. Not for use in diagnostic procedures.