Short-read amplicon sequencing targeting partial 16S rRNA variable regions (typically V3–V4 or V4–V5) provides genus-level classification at best — and in most microbiomes, 40–50% of sequences cannot be assigned to any known species. The fundamental limitation is not sequencing depth but read length: when the information-carrying region of a 16S rRNA gene spans approximately 1,500 bp, reading only 300–500 bp discards the phylogenetic signal contained in the remaining two-thirds of the sequence. Full-length 16S/18S/ITS amplicon sequencing overcomes this barrier by sequencing the complete gene in a single contiguous read, enabling species-level and, in many cases, strain-level taxonomic classification that is structurally impossible from short-read data alone.
CD Genomics provides full-length 16S (V1–V9), 18S (full-length SSU rRNA), and ITS (ITS1-5.8S-ITS2) amplicon sequencing on both PacBio Revio (HiFi circular consensus sequencing) and ONT PromethION (ultra-long single-molecule sequencing) platforms. Unlike sequencing providers that default to short-read approaches and offer long-read as a premium add-on, long-read sequencing is our core technology — every full-length amplicon project benefits from third-generation sequencing resolution from the start, on the platform best matched to the research question. We provide end-to-end service from primer design and long-range PCR optimization through library construction, platform-matched sequencing, and comprehensive bioinformatics analysis, delivering taxonomic profiles at a resolution that short-read methods cannot achieve.
The 16S rRNA gene contains nine variable regions (V1–V9) interspersed with conserved sequences that serve as universal PCR priming sites. Short-read amplicon sequencing targets 1–3 adjacent variable regions, capturing only 20–30% of the total phylogenetic information in the gene. This partial coverage imposes a hard ceiling on taxonomic resolution: at the genus level, most short-read studies achieve 80–85% classification; at the species level, classification rates drop to 40–55% regardless of sequencing depth, because the distinguishing nucleotides lie outside the sequenced fragments.
Full-length 16S sequencing reads all nine variable regions simultaneously, capturing the complete phylogenetic signal. The improvement is not marginal — it is qualitative. Species-level classification rates increase to 75–90% across diverse microbiome types, and for many genera, full-length sequences can distinguish closely related species that share >99% identity across V3–V4 alone. The same principle applies to the eukaryotic 18S SSU rRNA gene (~1,800 bp) and the fungal ITS region (~400–900 bp), where the internal transcribed spacer includes rapidly evolving regions that provide species-level discrimination within most fungal genera.
However, achieving full-length coverage is only half the equation. The platform used to sequence these long amplicons determines the accuracy, throughput, and cost structure of the project. CD Genomics offers two complementary long-read platforms, each optimized for different research priorities:
The choice between platforms — or the decision to use both in a complementary strategy — depends on the specific balance of accuracy, throughput, turnaround time, and budget that best fits each research project. Our project scientists provide platform-agnostic guidance to help you select the optimal approach.
Full-length amplicon sequencing is a targeted sequencing approach that amplifies and sequences complete ribosomal RNA genes or internal transcribed spacer regions in a single contiguous read, rather than sequencing short sub-fragments. For bacterial and archaeal communities, the target is the full-length 16S rRNA gene (~1,500 bp) encompassing all nine hypervariable regions (V1–V9). For eukaryotic microbial communities, the full-length 18S SSU rRNA gene (~1,800 bp) provides phylogenetic resolution across protists, microeukaryotes, and fungi. For fungal-specific profiling, the complete ITS region (ITS1-5.8S-ITS2, approximately 400–900 bp depending on species) captures the rapidly evolving spacer sequences that provide the highest taxonomic discrimination within the fungal kingdom.
Only long-read sequencing technologies — PacBio SMRT sequencing with circular consensus and ONT nanopore single-molecule sequencing — produce reads long enough to span these complete amplicons. Short-read sequencing platforms (Illumina, MGI) are physically limited to 2×300 bp paired-end reads (effectively ~500–550 bp after overlap merging), which can cover at most 2–3 adjacent variable regions of the 16S gene. This is not a protocol limitation — it is a fundamental read-length constraint that cannot be overcome by deeper sequencing or improved library preparation. Full-length amplicon sequencing removes this constraint, providing the complete phylogenetic information content of each marker gene for every sequenced amplicon.
Our service integrates both PacBio Revio and ONT PromethION platforms into a single service catalog, allowing researchers to select the optimal platform for each project's specific requirements — or to use both platforms in parallel for cross-validated, comprehensive community profiling.
Full-length 16S sequences routinely achieve 75–90% species-level classification rates across diverse microbiomes (gut, soil, marine, human-associated), compared to 40–55% from V3–V4 short-read data. The complete V1–V9 sequence captures the full phylogenetic signal needed to distinguish closely related species that are indistinguishable over partial gene fragments.
Simultaneous amplification of bacterial (16S), fungal (ITS), and microeukaryotic (18S) markers from the same DNA extracts enables cross-kingdom community analysis in a unified experimental design. This integrated approach reveals inter-kingdom interactions — bacteria–fungi, bacteria–protist — that are invisible when each kingdom is analyzed separately.
Full-length amplicon sequencing generates amplicon sequence variants (ASVs) with substantially higher phylogenetic resolution than short-read ASVs. The longer sequences improve the specificity of ASV clustering, reduce the incidence of chimeric ASVs, and enable more confident taxonomic placement of novel or uncultured lineages.
We match your project to the optimal platform: PacBio Revio for highest-accuracy strain-level resolution in small-to-medium cohort studies; ONT PromethION for ultra-high-throughput large cohort screens and real-time quality monitoring; or both for comprehensive cross-platform validation.
A single Revio SMRT Cell 8M processes 96–384 barcoded full-length 16S libraries in a 24-hour run. PromethION flow cells deliver comparable throughput with the added benefit of real-time data streaming. For studies requiring thousands of samples, we develop platform-optimized multiplexing strategies to maximize cost efficiency.
Our bioinformatics team implements platform-appropriate analysis pipelines: DADA2/QIIME2 for PacBio HiFi data, isONclust/LACA for ONT data, with cross-platform normalization when both data types are used in the same study. Deliverables include fully processed ASV tables, taxonomic assignments, diversity metrics, and publication-ready visualization.
Both platforms share an upstream workflow (DNA extraction, long-range PCR, library preparation, barcoding) but diverge in sequencing chemistry, data generation, and bioinformatics processing. Below we describe each platform's workflow independently, followed by a direct comparison to guide platform selection.
High-quality genomic DNA is extracted from the sample matrix (stool, soil, tissue, water filter, biofilm) using extraction protocols optimized for the sample type and target microbial groups. The full-length target region is amplified using platform-specific long-range PCR: 16S using universal primers 27F–1492R (or 27F–1540R for maximal V1–V9 coverage); 18S using primers targeting the complete SSU rRNA gene (~1,800 bp); ITS using primers ITS1–ITS4 covering the complete ITS1-5.8S-ITS2 region. PCR conditions are optimized to minimize amplification bias while maintaining amplicon length integrity. Amplicons are purified by AMPure bead cleanup to remove primers and short fragments, quantified, and assessed for size distribution by TapeStation or Fragment Analyzer.
Purified full-length amplicons are prepared for PacBio SMRT sequencing (see our PacBio SMRT sequencing technology page for platform details) using the SMRTbell Prep Kit 3.0. Amplicons are end-repaired and A-tailed, SMRTbell adapters are ligated, and libraries are size-selected using AMPure PB beads to retain full-length inserts. Barcoded libraries from different samples are pooled at equimolar ratios (up to 384 samples per SMRT Cell 8M), bound to polymerase, and loaded onto the Revio system. Each SMRT Cell 8M generates approximately 80–100 Gb of HiFi data in a 24-hour run. The circular consensus sequencing (CCS) algorithm reads each amplicon molecule multiple times (typically 10–20 passes per amplicon), generating a single high-accuracy (>Q30) consensus sequence per molecule. For full-length 16S amplicons (~1,500 bp), this translates to 200,000–500,000 HiFi reads per SMRT Cell at multiplexing levels of 96–384 samples, providing 2,000–10,000 reads per sample with per-read accuracy exceeding 99.9%. On-instrument basecalling produces CCS reads without additional bioinformatics infrastructure.
Full-length amplicons are prepared for nanopore sequencing (see our Oxford Nanopore sequencing technology page for platform details) using the ONT Ligation Sequencing Kit (SQK-LSK114) or native barcoding kit (SQK-NBD114.96) for multiplexed projects. Amplicons are end-prepped, dA-tailed, and ligated to sequencing adapters with attached motor proteins. Libraries are loaded onto R10.4.1 flow cells on the PromethION P48 or P24 platform. Each flow cell generates 100–290 Gb of data over a 72-hour sequencing run, with read N50 typically matching the amplicon length (e.g., ~1,500 bp for full-length 16S). Real-time basecalling (super-accurate or high-accuracy mode using Dorado) enables immediate data quality assessment and coverage monitoring during the run. For amplicon libraries of uniform length, PromethION flow cells can produce 10–50 million reads per flow cell, accommodating 96–384 barcoded samples with 10,000–50,000 reads per sample. Duplex basecalling (Q30+) is available for projects requiring the highest possible single-molecule accuracy from the nanopore platform.
Figure 1. Integrated full-length 16S/18S/ITS amplicon sequencing workflow. PCR amplicons are prepared in parallel for PacBio Revio SMRTbell libraries (circular consensus sequencing) and ONT PromethION libraries (single-molecule sequencing), followed by platform-matched sequencing and bioinformatics analysis.
PacBio HiFi data: CCS reads are demultiplexed by barcode, primer sequences are removed, and reads are quality-filtered. HiFi reads can be processed directly in DADA2 (via the "learnErrors" function adapted for HiFi error profiles) or QIIME2 with the q2-dada2 plugin to generate amplicon sequence variants (ASVs). Alternatively, de novo ASV clustering using isONclust provides reference-free ASV inference. Taxonomic assignment is performed against the SILVA 16S, GTDB, or UNITE ITS reference databases.
ONT PromethION data: Basecalled reads are demultiplexed, adapter-trimmed using Porechop, and quality-filtered (Q-score filtering). Because ONT reads have a distinct error profile (predominantly indels in homopolymer regions), standard DADA2 processing requires platform-specific adaptation. We deploy isONclust for reference-free consensus ASV inference, LACA (Long Amplicon Consensus Analysis) for denoising and error correction, or a modified DADA2 pipeline with ONT-specific error models. Taxonomic assignment follows the same reference databases used for HiFi data, with confidence scores adjusted for the platform-specific error profile.
| Analysis Feature | Standard Package | Advanced Package |
| Demultiplexing, primer removal, and read QC filtering | ✓ | ✓ |
| ASV/OTU clustering and taxonomic assignment (SILVA/GTDB/UNITE) | ✓ | ✓ |
| Alpha diversity (Shannon, Simpson, Chao1, observed ASVs) and rarefaction curves | ✓ | ✓ |
| Beta diversity (PCoA, NMDS, PERMANOVA) with taxonomic composition bar plots | ✓ | ✓ |
| Differential abundance analysis (DESeq2, ANCOM-BC, or LEfSe) | — | ✓ |
| Phylogenetic tree construction and UniFrac analysis | — | ✓ |
| Functional prediction (PICRUSt2, FAPROTAX, or custom pathway mapping) | — | ✓ |
| Cross-platform data integration and normalization (PacBio + ONT combined studies) | — | ✓ |
| Custom downstream analysis (network analysis, source tracking, longitudinal modeling) | — | ✓ |
| Publication-ready figures and summary report | ✓ Standard | ✓ Custom |
Both PacBio Revio and ONT PromethION deliver full-length amplicon sequencing with species-level resolution that short-read platforms cannot match. However, each platform has distinct performance characteristics that make it better suited for specific research contexts. The table below provides a side-by-side comparison to guide platform selection, drawing on published head-to-head evaluations including Biada et al. (2025, Frontiers in Microbiomes) and Hui et al. (2025, Gut Microbes).
| Feature | PacBio Revio | ONT PromethION | Dual-Platform Strategy |
| Read length (typical amplicon) | 15–25 kb (HiFi CCS); amplicon ~1,500 bp fully spanned | 20–100+ kb; amplicon matched read length with ultra-long capability | Full coverage on both platforms |
| Per-read accuracy (consensus) | ≥Q30 (>99.9%) — circular consensus, multiple passes per molecule | Q20+ (simplex); Q30+ (duplex) — R10.4.1 chemistry | HiFi accuracy + ONT duplex validation |
| Species-level classification rate (16S) | 63–75% (study-dependent; limited by reference database, not read quality) | 76% in recent studies — comparable when denoising is properly applied | Cross-validated species assignments |
| Error profile | Random errors — correctable by coverage; minimal homopolymer bias | Systematic indels in homopolymers — requires platform-specific error modeling | Error profiles are complementary |
| Throughput per run | ~90 Gb (SMRT Cell 8M); 200K–500K HiFi reads per cell | 100–290 Gb (flow cell); 10M–50M reads per flow cell | Maximum combined throughput |
| Multiplexing (typical 16S) | 96–384 samples per SMRT Cell | 96–384 samples per flow cell | Flexible per sample requirements |
| Run duration | ~24 hours | ~72 hours (real-time data available from 1 hour) | Dependent on primary platform |
| Basecalling infrastructure | On-instrument — no additional hardware required | GPU recommended for real-time Dorado basecalling | Handled by our bioinformatics team |
| Best suited for | Highest-accuracy strain-level resolution, rare variant detection, DADA2-compatible ASV inference | Large cohort screens, native modification detection, real-time quality monitoring, ultra-high read depth | Comprehensive cross-platform studies requiring both accuracy and depth |
In practice, the most important consideration is how the data will be analyzed. PacBio HiFi data is directly compatible with established short-read ASV pipelines (DADA2, QIIME2) because its error profile matches the assumptions of these tools. ONT data requires platform-aware processing (isONclust, LACA, or ONT-modified DADA2) but delivers comparable biological conclusions when properly analyzed. For projects where both accuracy and read depth are critical, our dual-platform strategy provides the most comprehensive solution, integrating PacBio HiFi accuracy with ONT ultra-deep coverage in a unified analysis framework.
| Category | Requirement | Notes |
| Sample type | gDNA (stool, soil, tissue, water filter, biofilm, swab); or purified amplicon libraries | gDNA extraction service available for challenging sample types |
| Minimum input (gDNA) | 10–100 ng (sufficient for 16S/18S/ITS long-range PCR) | Lower inputs possible for high-bacterial-biomass samples; 100 ng recommended for low-biomass samples |
| Minimum input (amplicon library) | ≥200 ng purified amplicon (quantified by Qubit) | Amplicon size verified by TapeStation or Fragment Analyzer before library preparation |
| DNA quality | A260/280 ≥ 1.8; minimal humic acid or polyphenol contamination (for soil/sediment samples) | Additional purification steps available for challenging environmental samples |
| Recommended reads per sample | 2,000–5,000 (PacBio HiFi); 5,000–20,000 (ONT PromethION) | Higher read depth recommended for low-biomass or high-diversity samples |
| Multiplexing | 96–384 samples per SMRT Cell or flow cell | Custom barcoding strategies available for projects requiring >384 samples |
| Shipping | Overnight on dry ice (gDNA) or room temperature (amplicon libraries) | See our sample submission guidelines for detailed instructions |
Long-read specialist — not a short-read provider with a long-read side service.
CD Genomics is a long-read sequencing-focused CRO. Our full-length amplicon sequencing service is built around PacBio Revio and ONT PromethION platforms, not around NGS with long-read as a premium add-on. This means every project receives platform-optimized experimental design, library preparation, and bioinformatics from a team that thinks in long reads, not short reads.
True dual-platform flexibility with platform-agnostic guidance.
We operate both PacBio Revio and ONT PromethION platforms in-house, with dedicated teams for each platform's library chemistry and sequencing workflow. We help you select the platform — or the platform combination — that best matches your research priorities, without pushing you toward one platform over another based on availability or margin.
Platform-appropriate bioinformatics — we know the difference matters.
PacBio HiFi and ONT data require fundamentally different bioinformatics approaches. We do not apply a one-size-fits-all pipeline. Our analysis workflows are platform-matched: DADA2/QIIME2 for HiFi data, isONclust/LACA for ONT data, with validated cross-platform normalization when both data types are integrated in a single study. Every deliverable includes platform-specific quality metrics so you know exactly what each dataset represents.
Proven track record across diverse microbiome applications.
We have delivered full-length amplicon sequencing projects spanning human gut, soil, marine, plant-associated, industrial fermentation, and clinical microbiomes, with sample processing volumes from pilot-scale (dozens of samples) to population-scale (thousands of samples). Our published case study collaborations demonstrate the species-level resolution and biological insights that full-length amplicon sequencing — on the right platform — can deliver.
Hui Y, Nielsen DS, Krych L. De novo clustering of long-read amplicons improves phylogenetic insight into microbiome data. Gut Microbes. 2025;17(1):2516703. (CC BY 4.0)
The promise of long-read full-length 16S rRNA amplicon sequencing — species-level taxonomic resolution — has been constrained by the lack of robust, platform-appropriate bioinformatics methods for processing long-read amplicon data. While short-read amplicon data has mature analysis ecosystems (DADA2, QIIME2, USEARCH), long-read data generated on ONT and PacBio platforms presents distinct error profiles that require de novo clustering approaches. Hui and colleagues developed and evaluated the LACA (Long Amplicon Consensus Analysis) workflow — a de novo clustering pipeline designed specifically for long-read amplicon data — benchmarking it across multiple ONT chemistries, PacBio CCS data, and diverse microbiome sample types including synthetic mock communities, human vaginal microbiomes, and bacterial isolates.
The study evaluated full-length 16S rRNA amplicon sequencing data generated on multiple technology configurations: ONT MinION R9.4.1, R10.3, and R10.4.1 flow cells with simplex and duplex basecalling, and PacBio CCS (circular consensus sequencing). Samples included a synthetic 7-species bacterial mock community, 10 human vaginal microbiome samples, and whole-genome sequencing data from bacterial isolates for ground-truth validation. The LACA workflow employed HDBSCAN density-based clustering of k-mer frequencies to generate de novo consensus ASVs without reference-based error correction, and was benchmarked against isONclust, DADA2 (with long-read modifications), and traditional OTU clustering at 97% and 99% identity thresholds.
Figure 2. Performance of long-read amplicon clustering approaches across ONT and PacBio platforms. (A) Error rates for different sequencing chemistries — ONT R9.4.1 (<1%), R10.3 (<1%), R10.4.1 (<0.2%), ONT Duplex (<0.1%), and PacBio CCS (<0.1%). (B) Species-level taxonomic assignment accuracy for LACA and alternative clustering methods. From Hui et al. (2025, Gut Microbes, CC BY 4.0).
This study provides critical validation for two conclusions directly relevant to full-length amplicon sequencing service design. First, both ONT and PacBio platforms — when paired with appropriate bioinformatics — deliver accurate species-level taxonomic resolution from full-length 16S amplicons, with the accuracy gap between platforms narrowing substantially with R10.4.1 chemistry and duplex basecalling. Second, the choice of bioinformatics method matters at least as much as the sequencing platform: reference-free de novo clustering methods (LACA, isONclust) provide better species discrimination than fixed-threshold OTU clustering, particularly for closely related species. These findings directly inform our dual-platform service design — we match both the sequencing platform AND the bioinformatics approach to the specific requirements of each project, rather than applying a single pipeline to all data regardless of platform.
Full-length 16S sequencing routinely achieves species-level classification rates of 75–90%, compared to 40–55% from V3–V4 short-read data. The improvement is most pronounced for genera containing closely related species with distinct ecological roles or pathogenic potential — Bacteroides, Clostridium, Lactobacillus, Streptococcus, and Bacillus are among the genera where full-length sequences routinely resolve multiple species that are indistinguishable from V3–V4 alone. However, species-level classification is ultimately limited by reference database completeness, not sequencing technology — for under-studied environments where reference genomes are sparse, full-length sequences may still assign to genus level due to the absence of closely related reference sequences.
Published head-to-head comparisons (Biada et al. 2025, Hui et al. 2025) show that both platforms achieve comparable species-level classification rates when platform-appropriate bioinformatics is applied — typically 63–76% for PacBio HiFi and a similar range for ONT with modern R10.4.1 chemistry. PacBio HiFi reads have intrinsically higher per-read accuracy (Q30+) and integrate directly with established DADA2-based pipelines. ONT reads achieve comparable biological conclusions when processed with platform-specific denoising tools (isONclust, LACA). The practical difference between platforms is smaller than the difference between either long-read platform and short-read V3–V4 sequencing. The choice between PacBio and ONT should be driven by throughput requirements, project scale, and bioinformatics preferences rather than by a significant difference in taxonomic resolution.
Yes — we routinely design multiplexed panels co-amplifying bacterial 16S, eukaryotic 18S, and fungal ITS markers from the same DNA samples in a single reaction or parallel reactions, followed by pooled library preparation and sequencing on a single SMRT Cell or flow cell. The different amplicon sizes (~1,500 bp for 16S, ~1,800 bp for 18S, ~400–900 bp for ITS) are compatible with both PacBio Revio and ONT PromethION platforms. Each amplicon type is identified by amplicon-specific barcoding or size-based separation during bioinformatics. This multi-kingdom approach enables integrated analysis of bacterial, fungal, and microeukaryotic communities from the same samples without additional sequencing costs.
Yes — this is one of the key practical advantages of PacBio HiFi data for full-length amplicon sequencing. Because HiFi consensus reads have a random error profile (rather than the systematic error profile of ONT reads) and achieve >Q30 accuracy, they are directly compatible with DADA2's error learning algorithm and QIIME2's q2-dada2 plugin. Simply provide the CCS FASTA/Q files as input, and the standard DADA2 workflow processes them correctly. For ONT data, DADA2 can be used with platform-specific error models or replaced by tools designed for ONT error profiles (isONclust, LACA). Our bioinformatics team implements the appropriate pipeline for each data type and provides guidance on analysis tool selection during project design.
Recommended sequencing depth depends on the expected diversity of the microbial community and the taxonomic resolution required. For moderate-complexity communities (human gut, vaginal, skin microbiomes), 2,000–5,000 PacBio HiFi reads per sample or 5,000–10,000 ONT reads per sample are typically sufficient for comprehensive species-level profiling. For high-diversity communities (soil, sediment, marine), 5,000–10,000 HiFi reads or 10,000–20,000 ONT reads per sample are recommended. These depths are substantially lower than the 50,000–100,000 reads typically required for short-read V3–V4 sequencing because each full-length read carries more phylogenetic information than a short-read fragment. Our multiplexing strategies are designed to deliver these coverage depths cost-effectively across 96–384 samples per sequencing run.
1. Per-sample ASV/OTU abundance table with taxonomic assignments from Kingdom to Species level (SILVA 16S, GTDB, or UNITE ITS reference databases)
2. Alpha diversity metrics and rarefaction curves with statistical comparisons between experimental groups
3. Beta diversity ordination plots (PCoA, NMDS) with PERMANOVA significance testing and taxonomic composition bar plots at multiple taxonomic ranks
4. Differential abundance analysis results with effect sizes, adjusted p-values, and volcano/heatmap visualization
5. Optional: functional prediction profiles (PICRUSt2), phylogenetic trees, and cross-kingdom correlation networks
References
For research use only. Not for use in diagnostic procedures.