PacBio Sequencing Data Analysis — Turn Revio HiFi Reads into Publication-Ready Results

PacBio Sequencing Data Analysis — Turn Revio HiFi Reads into Publication-Ready Results

PacBio HiFi sequencing data analysis pipeline for Revio reads, covering variant calling, assembly, and methylation

PacBio HiFi sequencing, generated on the current Revio system, produces reads with a median length of 15-20 kb and consensus accuracy above 99.9%, giving researchers unbiased, near-complete coverage of the genome, transcriptome, or epigenome in a single dataset. Raw HiFi reads only become useful once they are aligned, called, assembled, and annotated—and PacBio's toolchain has evolved substantially since the Sequel-generation era. CD Genomics' PacBio sequencing data analysis service runs the current, community-standard toolchain—pbmm2, DeepVariant, pbsv, HiPhase, TRGT, HiFiCNV, hifiasm, and the Iso-Seq analysis suite—to turn Revio HiFi data into variant calls, assemblies, isoforms, and methylation maps ready for interpretation or publication.

Whether your HiFi data comes from our own sequencing service or was generated elsewhere, our dedicated bioinformatics team builds a workflow around your specific question—human, plant, animal, or microbial—rather than running a one-size-fits-all pipeline.

Why researchers choose our PacBio data analysis service

Introduction

Long-read HiFi sequencing on the PacBio Revio system generates far more than raw reads—within a single dataset lie SNVs, indels, structural variants, tandem repeat expansions, copy number changes, haplotype phase, full-length transcript isoforms, and CpG methylation status. Extracting each of these signals reliably requires the right tool at each step, applied with the right parameters for your organism and question. Tool choices that were standard in the Sequel II era—built around continuous long reads (CLR) rather than HiFi consensus reads—no longer reflect current best practice. CD Genomics keeps its PacBio analysis pipelines current with the tools PacBio and the long-read community actively maintain and benchmark today.

What Is PacBio Sequencing Data Analysis?

PacBio sequencing data analysis is the set of computational steps that convert raw HiFi reads into interpretable biological results. Starting from unaligned BAM files produced by the Revio system, reads are demultiplexed, aligned to a reference (or assembled de novo when no reference exists), and passed through variant calling, phasing, structural variant, tandem repeat, and—where relevant—transcript or methylation analysis modules.

Because HiFi reads already carry near-perfect per-base accuracy from the sequencer, this analysis focuses on structural and biological interpretation rather than error correction, which was a major burden with older continuous long-read (CLR) data.

Our team runs this analysis using PacBio SMRT sequencing data as the primary input, and can integrate results alongside Oxford Nanopore-derived datasets for projects that combine both platforms.

Data Sources We Analyze

Our pipelines are built around HiFi data from the current PacBio instrument generation, with support for legacy datasets where needed.

PacBio Revio HiFi Data (Primary)

The Revio system, running SPRQ-Nx chemistry, is our primary data source. Revio HiFi reads carry a median consensus accuracy of Q30 or better, with on-instrument 5mC (and 5hmC) methylation calling already embedded in the aligned BAM. We align this data with pbmm2, a minimap2 wrapper purpose-built for HiFi reads.

Legacy Sequel/Sequel II Datasets

For customers with existing Sequel or Sequel II HiFi (CCS) data, we can still run the modern variant-calling and assembly toolchain, though we recommend Revio for new sequencing given its higher throughput and lower per-genome cost.

Data-analysis-only engagements are welcome—send us your existing HiFi BAM or FASTQ files. If you also need new sequencing, we offer PacBio library construction and Revio sequencing through our PacBio Pre-Made Library Sequencing service.

Key Advantages of Our Analysis Service

Scientific Advantages

  • Current, benchmarked toolchain

pbmm2, DeepVariant, pbsv, HiPhase, TRGT, and HiFiCNV form our default WGS variant workflow, matching the pipeline PacBio itself recommends for Revio data.

  • Haplotype-resolved output

HiPhase jointly phases small variants, structural variants, and tandem repeats using HiFi read-backed phasing, giving allele-specific results without a separate assay.

  • Assembly without a reference

hifiasm and verkko deliver telomere-to-telomere-quality, haplotype-resolved de novo assemblies directly from HiFi reads, with optional Hi-C or trio-binning integration.

  • Multi-omic from one dataset

Full-length isoform identification (Iso-Seq) and CpG methylation profiling can both be run on the same HiFi reads used for variant calling, maximizing the value of each SMRT Cell.

Business & Project Advantages

  • Flexible entry point

Bring your own HiFi data, or combine sequencing and analysis as a single project—both are supported.

  • Organism-agnostic pipelines

The same core toolchain scales from human clinical-research samples to plant, animal, and microbial genomes, with organism-specific reference and parameter tuning.

  • Reproducible, containerized workflows

Our pipelines run in containerized, version-controlled environments, so results are reproducible and auditable for publication.

  • Clear, interpretable reports

Deliverables include annotated VCFs, assembly statistics, isoform tables, and methylation tracks, alongside a plain-language summary report.

Applications

Whole-Genome Variant Detection

De Novo Genome Assembly

Full-Length Transcriptome Analysis (Iso-Seq)

Epigenetic and Complex-Population Analysis

Analysis Workflow - How It Works

1. Pre-Processing

Unaligned HiFi BAM files from the Revio system are demultiplexed if needed, then aligned to a reference genome with pbmm2, or routed directly to assembly if no reference exists. Reads are sorted and indexed for downstream tools.

2. Variant Calling and Phasing

SNVs and small indels are called with DeepVariant; structural variants with pbsv; tandem repeats with TRGT; and copy number variants with HiFiCNV. HiPhase then jointly phases all variant types using HiFi read-backed information.

3. Assembly or Transcript/Methylation Analysis

For de novo projects, hifiasm or verkko builds haplotype-resolved contigs directly from HiFi reads. For transcriptomic projects, full-length reads are classified into isoforms. For epigenetic projects, on-instrument 5mC/5hmC calls are extracted and mapped to genomic context.

4. Annotation and Reporting

  • Variants are annotated against relevant databases and filtered by quality and population frequency where applicable
  • Assembly statistics (N50, BUSCO completeness, phasing quality) are calculated
  • Isoform tables and methylation tracks are generated for visualization in IGV or similar tools
  • A summary report and a MultiQC-style run report are compiled for every project

PacBio HiFi data analysis workflow from Revio reads through variant calling, assembly, and annotationWorkflow of PacBio HiFi data analysis, from raw Revio reads through alignment or assembly, variant calling and phasing, to annotated, publication-ready output.

Analysis Modules and Tools

Analysis Module Current Toolchain Detects Typical Applications
Whole-genome variant calling pbmm2, DeepVariant, pbsv, TRGT, HiFiCNV, HiPhase SNVs, indels, SVs, tandem repeats, CNVs, phased haplotypes Human, plant, animal, and microbial genome research
De novo assembly hifiasm, verkko Haplotype-resolved, near-gapless genome assemblies Reference genome construction, pan-genome projects
Full-length RNA analysis (Iso-Seq) SMRT Link Iso-Seq pipeline, isoform classification and quantification tools Full-length isoforms, novel transcripts, fusion events Transcriptome annotation, alternative splicing studies
Epigenetics and base modification On-instrument 5mC/5hmC calling, SMRT Link kinetics tools CpG methylation, DNA/RNA base modification, motif analysis Allele-specific methylation, epigenetic regulation studies
Complex population analysis pbmm2, DeepVariant, pbsv, minor-variant calling workflows De novo assembly, minor variant and modification detection Bacterial, fungal, and viral population studies

Choosing the Right Analysis Module

Most projects combine more than one module. The table below summarizes what each core module adds, to help you scope your project.

Module Reference Genome Needed? Best For Not Ideal For
WGS variant pipeline Yes SNV/indel/SV/CNV detection, phasing, repeat expansion genotyping Organisms without a usable reference genome
De novo assembly (hifiasm/verkko) No New reference genomes, non-model organisms, pan-genome studies Quick variant screening against a known reference
Iso-Seq transcript analysis Recommended, not strictly required Full-length isoform discovery, alternative splicing, fusion detection DNA-level structural variant detection
Methylation/kinetics analysis Yes Allele-specific methylation, epigenetic regulation Samples without matched genomic context

How to interpret this comparison

  • Start with the WGS variant pipeline when a reference genome is available and your priority is comprehensive variant detection with phasing.
  • Choose de novo assembly when no adequate reference exists, or when you need a new haplotype-resolved reference for your species.
  • Add Iso-Seq or methylation modules on top of either pipeline when your project also requires transcript-level or epigenetic information from the same samples.

Data Input Requirements

Category Requirement Notes
File format Unaligned HiFi BAM (preferred) or FASTQ Native PacBio BAM preserves kinetics tags needed for methylation calling
Instrument generation PacBio Revio (preferred); Sequel/Sequel II HiFi (CCS) also supported CLR-only data requires additional preprocessing and is not recommended for new projects
Minimum coverage - WGS variant calling ≥ 15-20× for germline SNVs/indels; ≥ 30× recommended for full SV/CNV/phasing confidence Coverage needs scale with genome size and heterozygosity
Minimum coverage - de novo assembly ≥ 20× HiFi coverage, higher for highly heterozygous or polyploid genomes Hi-C or trio data further improves haplotype resolution
Reference genome Provided by customer, or selected jointly from public databases Required for alignment-based modules; not required for de novo assembly
Data transfer Secure cloud transfer or physical hard drive We can also receive data directly if sequencing was performed by CD Genomics

Why Choose CD Genomics

Current, Community-Aligned Toolchain

We run the same core tools—pbmm2, DeepVariant, pbsv, HiPhase, TRGT, HiFiCNV, hifiasm—that define current best practice for Revio HiFi data analysis.

Flexible Engagement

Analyze data you already have, or combine it with our Revio sequencing and library construction services for an end-to-end project.

Multi-Omic Expertise

From variant calling to de novo assembly, Iso-Seq, and methylation analysis, our bioinformaticians work across the full range of HiFi applications.

Organism-Agnostic Experience

Human, plant, animal, and microbial genome projects are all supported, with reference and parameter choices tailored to your organism. Structural variant work often pairs well with our dedicated human genome structural variation detection service.

Transparent, Reproducible Reporting

Every project is delivered with containerized, version-tracked workflows and a clear summary report suitable for methods sections and internal review.

Case Study: A Reproducible Pipeline for PacBio WGS and Repeat Expansion Analysis

Jain, T., Clelland, C. nf-core/pacvar: a pipeline for analyzing long-read PacBio whole genome and repeat expansion sequencing data. Bioinformatics 41(4), btaf116 (2025).

1. Background

PacBio long-read sequencing enables both whole-genome variant detection and targeted characterization of complex repeat-expansion loci linked to neurodegenerative disease, through PacBio's PureTarget panel. Existing toolkits such as SMRT Link and PacBio's HiFi-human-WGS-WDL pipeline each carried infrastructure limitations—SMRT Link could not run natively on macOS, and the WDL pipeline lacked compatibility with certain HPC schedulers and cloud platforms. There was a clear need for a single, portable pipeline covering both use cases.

2. Methods

The authors built nf-core/pacvar, a Nextflow pipeline with three main components:

The pipeline was built to run identically across HPC clusters with different schedulers, cloud platforms, and local systems, using Docker, Singularity, Apptainer, or conda containers.

3. Results

nf-core/pacvar pipeline structure for PacBio HiFi whole genome and repeat expansion analysisThe nf-core/pacvar pipeline structure, comprising pre-processing, variant calling and phasing, and repeat expansion characterization, with the specific tool used at each step.

Key Findings

4. Conclusions

This study demonstrates that a modular, containerized pipeline built on current PacBio HiFi tools can deliver both comprehensive WGS variant detection and specialized repeat-expansion characterization in a single, portable workflow. Importantly:

FAQs

Demo

1. Phased Small Variant, Structural Variant, and Tandem Repeat Overview

2. Haplotype-Resolved Assembly Contiguity Summary

3. Repeat Expansion Genotyping Waterfall Plot

demo figures for PacBio HiFi variant phasing, assembly contiguity, and repeat expansion genotyping

References

  1. Jain, T., Clelland, C. nf-core/pacvar: a pipeline for analyzing long-read PacBio whole genome and repeat expansion sequencing data. Bioinformatics. 41(4), btaf116 (2025).
  2. Wenger, A.M., Peluso, P., Rowell, W.J. et al. Accurate circular consensus long-read sequencing improves variant detection and assembly of a human genome. Nat Biotechnol. 37, 1155-1162 (2019).
  3. Cheng, H., Concepcion, G.T., Feng, X. et al. Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm. Nat Methods. 18, 170-175 (2021).
  4. Holt, J.M., Saunders, C.T., Rowell, W.J. et al. HiPhase: jointly phasing small, structural, and tandem repeat variants from HiFi sequencing. Bioinformatics. 40, btae042 (2024).

For Research Use Only. Not for use in diagnostic procedures.

Get Your Instant Quote