Comprehensive human genome sequencing on PacBio HiFi (Revio / Sequel II) and Oxford Nanopore PromethION — delivering genome-wide structural variant detection, telomere-to-telomere genome analysis, haplotype phasing, and integrated epigenome profiling for rare disease research, cancer genomics, and population-scale studies
CD Genomics provides end-to-end human whole-genome sequencing services on PacBio HiFi (Revio/Sequel II) and Oxford Nanopore PromethION platforms — from HMW DNA extraction and library preparation through multi-caller structural variant detection, tandem repeat analysis, mobile element insertion discovery, de novo assembly, haplotype phasing, and pharmacogenomic annotation — delivering genome-wide SV detection at nucleotide resolution from a single sequencing experiment.
Human whole-genome sequencing (WGS) by long-read technology represents a fundamental advance in our ability to detect the full spectrum of human genetic variation. While short-read sequencing has been the dominant approach for the past two decades, its inherent read-length limitation (150–300 bp) means that approximately 5–10% of the human genome — including most structurally complex regions — remains inadequately resolved. Long-read sequencing platforms from Pacific Biosciences (HiFi on Revio / Sequel II) and Oxford Nanopore (PromethION) produce reads of 15–100+ kilobases, providing the single-molecule resolution needed to characterize structural variants (SVs), tandem repeats, mobile element insertions, and complex rearrangements that are systematically missed or mischaracterized by short-read approaches.
The impact of this technological shift is substantial. Large-scale studies such as the All of Us research program have demonstrated that long-read WGS at moderate coverage (8–10× HiFi) detects over twice as many disease-associated SVs as short-read WGS at 30× coverage, revealing previously hidden genetic contributions to rare and common diseases. Long-read WGS also provides haplotype-resolved genome data, direct detection of DNA base modifications without chemical pretreatment, and the ability to produce telomere-to-telomere genome assemblies — delivering a comprehensive genomic picture from a single sequencing experiment.
At CD Genomics, we provide human whole-genome sequencing services across both major long-read platforms — PacBio HiFi (Revio, Sequel II) and Oxford Nanopore PromethION — with coverage tailored to the specific requirements of each project. Whether your research involves rare disease variant discovery, somatic SV detection in cancer, population-scale genome characterization, or pharmacogenomic profiling, our team will design and execute a sequencing and analysis strategy optimised for your research objectives.
Human whole-genome sequencing aims to determine the complete DNA sequence of a human genome, covering all 22 autosomes, the X and Y chromosomes, and the mitochondrial genome. The goal is not simply to produce a sequence, but to generate a comprehensive and accurate catalog of all genetic variation present in the sample — from single-nucleotide variants (SNVs) and small insertions/deletions (InDels) to large structural variants, repeat expansions, and mobile element insertions — and to place this variation in its haplotype context.
The limitations of short-read WGS for this task are now well documented. The 150–300 bp reads produced by Illumina sequencers are shorter than most human repetitive elements and structural variant breakpoints, making it impossible to uniquely map reads across these regions. Studies comparing long-read and short-read WGS in the same individuals have consistently shown that short-read data misses 50–70% of SVs, particularly insertions (which are the most common SV type in the human genome), tandem repeat expansions, and complex rearrangements involving segmental duplications. For clinical applications, this means that patients with SV-mediated genetic diseases may remain undiagnosed after short-read WGS, and potentially actionable pathogenic variants remain hidden in uncharacterized genomic regions.
Long-read WGS overcomes this limitation fundamentally. A single PacBio HiFi read of 15–25 kb or an ONT read of 50–100+ kb can span an entire SV breakpoint junction, resolve a tandem repeat array, or bridge a segmental duplication to place structural variation in its correct genomic context. With coverage as low as 10–15× HiFi, long-read WGS achieves comprehensive SV detection across the genome. At higher coverage (20–30×), it enables de novo assembly of individual genomes, haplotype-level phasing, and complete characterization of the 5–10% of the genome that is inaccessible to short reads.
At CD Genomics, our human WGS service is designed to deliver the full value of long-read sequencing — whether that means comprehensive SV discovery at moderate coverage for cohort studies, high-coverage de novo assembly for reference-quality genomes, or integrated genome and epigenome analysis for multi-omics research projects.
Long-read WGS detects all major SV classes — deletions, duplications, inversions, translocations, mobile element insertions (Alu, LINE1, SVA), and tandem repeat expansions — at nucleotide resolution. Studies consistently report 2–3× more SVs detected by long reads compared to short reads in the same genome, with the majority of novel SVs being insertions and complex rearrangements in repetitive or duplicated regions.
PacBio HiFi reads with hifiasm enable production of haplotype-resolved assemblies that separate maternal and paternal chromosomes. This phased information is critical for understanding compound heterozygosity, parent-of-origin effects, allele-specific expression, and meiotic recombination — information that is computationally inferred at best from short-read data.
Both PacBio HiFi (kinetics-based 5mC, 6mA detection) and ONT (direct electrical current-based 5mC, 5hmC, 6mA detection) provide genome-wide methylation data from the same sequencing run used for variant detection — enabling integrated genetic and epigenetic analysis with no additional sample preparation or sequencing cost.
Moderate coverage long-read WGS (10–15× HiFi) provides comprehensive SV discovery at a cost significantly lower than high-coverage short-read WGS when accounting for the additional value of phased data, methylation information, and access to previously hidden variant classes — delivering more biological insight per sequencing dollar.
From 8–10× HiFi for population-scale screening (All of Us-style studies) to 30× HiFi for de novo assembly and comprehensive variant benchmarking, we recommend and deliver coverage tailored to your specific research objectives and budget.
Our bioinformatics pipelines cover the complete analysis spectrum: read alignment and QC, multi-caller SV detection (alignment-based and assembly-based), SNV/InDel calling, repeat expansion analysis, methylation calling, pharmacogenomic annotation, and custom variant filtration — delivered with comprehensive reporting for publication and downstream analysis.
Beyond rare disease diagnosis, cancer genomics, and population studies, long-read WGS is increasingly applied in resolving somatic mosaicism and age-related clonal hematopoiesis, where detection of structural variants present at low variant allele fractions (5–10%) in a subset of cells requires the single-molecule resolution that only long-read sequencing provides. Mosaic SVs — including mobile element insertions, complex rearrangements, and large deletions — arise in repetitive and duplicated genomic regions that are systematically inaccessible to short-read SV detection, making long-read WGS the only current approach capable of characterizing the full spectrum of somatic structural variation in aging and diseased tissues.
Each human WGS project begins with a consultation to determine the research objectives — germline variant discovery, somatic mutation detection, or population screening — from which we recommend the optimal sequencing platform, coverage target, and analysis pipeline. Sample quality is assessed by Qubit fluorometry, UV spectrophotometry, and pulsed-field gel electrophoresis to confirm HMW DNA integrity (fragments ≥ 30 kb for PacBio; ≥ 20 kb for ONT).
Platform-specific libraries are prepared according to the selected strategy. For PacBio HiFi, SMRTbell libraries with 15–20 kb inserts are prepared using standard or low-input protocols depending on available DNA quantity, with barcoded overhang adapters for multiplexing. For Oxford Nanopore, native ligation libraries (SQK-LSK114) are prepared with optional ultra-long read enrichment for spanning the most complex genomic regions. For hybrid strategies, both library types are prepared from the same DNA sample.
Sequencing is performed to the target coverage agreed during project design. Coverage recommendations are: 8–15× HiFi for population screening and SV discovery; 15–25× HiFi or ONT for comprehensive variant detection including SNVs and InDels; 25–30× HiFi for de novo assembly and the highest quality variant benchmarking. Sequencing is performed on PacBio Revio or Sequel II systems (HiFi CCS mode) and ONT PromethION (R10.4.1 flow cells, Dorado SUP basecalling).
Figure 1. Complete human whole-genome sequencing workflow — from sample QC and HMW DNA extraction through PacBio HiFi or ONT long-read sequencing, variant calling, and comprehensive genome analysis.
Raw sequencing data is processed through our bioinformatics pipeline: basecalling (Dorado SUP for ONT, SMRT Link for PacBio), read quality filtering and adapter removal, alignment to the human reference genome (GRCh38 or T2T-CHM13) using minimap2 or pbmm2, and initial quality metrics (coverage depth, uniformity, insert size distribution, mapping rate).
Downstream analysis includes multi-caller SV detection using both alignment-based (Sniffles2, cuteSV, SVIM) and assembly-based (pav, dipcall) approaches for comprehensive SV discovery; SNV/InDel calling with DeepVariant or Clair3; tandem repeat analysis (TRGT, Straglr); mobile element insertion detection (xTea, MELT); DNA methylation calling; and pharmacogenomic annotation. Results are compiled into a comprehensive project report with publication-ready figures and data files.
Our bioinformatics pipeline for human WGS data supports comprehensive genome analysis from raw read processing through multi-platform variant detection, genome assembly, and advanced interpretation — organized across two service tiers to accommodate different research objectives and budgets.
| Analysis Feature | Basic Package | Advanced Package |
| Read QC & alignment | ✓ Dorado/SMRT Link basecalling; quality filtering; minimap2/pbmm2 alignment to GRCh38; coverage metrics | ✓ + Alignment to T2T-CHM13; multi-reference comparison; coverage depth correction for GC bias |
| Structural variant detection | ✓ Sniffles2 or cuteSV; SV type classification; VCF output with allele frequencies | ✓ + Multi-caller (Sniffles2 + cuteSV + SVIM + pbsv); assembly-based SV detection; complex SV resolution; tandem repeat analysis (TRGT) |
| SNV & InDel calling | ✓ DeepVariant or Clair3; standard VCF output; variant quality score recalibration | ✓ + Ensemble calling; trio/phased family analysis; multi-sample joint calling for cohort studies |
| Haplotype phasing | — | ✓ HiFi-based phasing (hifiasm/whatshap); SNP-SV combined phasing; haplotype-resolved assembly output |
| De novo genome assembly | — | ✓ hifiasm diploid assembly; contig-level assembly quality metrics (N50, BUSCO, Merqury) |
| Mobile element insertion detection | ✓ xTea or MELT; Alu/LINE1/SVA classification; insertion allele frequency estimation | ✓ + MEI breakpoint sequence resolution; reference mobile element absence confirmation |
| Tandem repeat analysis | ✓ ExpansionHunter or Straglr; known pathogenic repeat loci screening | ✓ + TRGT repeat motif analysis; full repeat array sequence assembly; de novo repeat discovery |
| Copy number variation | ✓ CNVnator or QDNAseq; read-depth based CNV detection | ✓ + Multi-caller integration; allele-specific CNV; mosaic CNV detection at low variant allele fraction |
| DNA methylation analysis | — | ✓ 5mC/5hmC detection from HiFi kinetics or ONT raw signal; methylation frequency per CpG; differentially methylated region analysis |
| Pharmacogenomic annotation | — | ✓ PharmGKB star allele calling; CYP/UGT locus structural analysis; phenotype prediction report |
| Custom reporting & visualization | ✓ Standard project report PDF with coverage metrics, variant summary tables, and SV classification | ✓ IGV browser session files; publication-ready SV validation figures; interactive circos genome maps; VCF/BCF data files; NCBI BioProject submission support |
The optimal sequencing strategy for human WGS depends on the variant classes of primary interest, required detection sensitivity, coverage depth, and project budget. We provide platform-neutral recommendations based on the specific requirements of each research project.
| Feature | PacBio HiFi | Oxford Nanopore | Hybrid (HiFi + ONT) |
| Read length | 15–25 kb (CCS consensus) | 20–100+ kb (native) | Multi-platform combination |
| Per-base accuracy | ★★★★★ (Q30+; >99.9%) | ★★★☆☆ (Q14–Q20 raw; Q30+ with duplex) | ★★★★★ (HiFi-polished) |
| SV detection sensitivity | ★★★★☆ (excellent at ≥15×) | ★★★★☆ (excellent at ≥20×) | ★★★★★ (comprehensive) |
| SNV/InDel accuracy | ★★★★★ (comparable to short-read) | ★★★☆☆ (improving with duplex) | ★★★★★ (cross-validated) |
| Haplotype phasing | ★★★★★ (hifiasm phasing) | ★★☆☆☆ (limited without HiFi) | ★★★★★ (HiFi phasing + ONT contiguity) |
| De novo assembly | ★★★★★ (high-quality diploid assembly) | ★★★★☆ (ultra-long contiguity) | ★★★★★ (best quality) |
| Repeat expansion analysis | ★★★☆☆ (15–25 kb reads) | ★★★★★ (50–100 kb spans full arrays) | ★★★★★ (both accuracy and span) |
| DNA methylation detection | ★★★☆☆ (5mC, 6mA kinetics) | ★★★★★ (5mC, 5hmC, 6mA, 4mC native signal) | ★★★★★ (multi-modification) |
| Coverage for SV discovery | 10–15× | 15–25× | 10× HiFi + 15× ONT |
| Coverage for assembly | 25–30× | 30–60× | 20× HiFi + 20× ONT |
| Per-genome cost (multiplexed) | $$$ (moderate) | $ (lowest) | $$$ (highest, best quality) |
| Best suited for | Precision SNV/SV calling; haplotype phasing; de novo assembly; clinical research requiring high accuracy | Ultra-long reads for repeat expansion; large cohort screening; cost-sensitive projects; methylation-focused studies | Comprehensive variant discovery; T2T diploid assembly; integrated genome + epigenome projects |
For most human WGS projects, we recommend a primary PacBio HiFi strategy (15–20×) as the optimal balance of variant detection sensitivity, base accuracy, haplotype phasing, and cost. For projects prioritising ultra-long read information (repeat expansions, complex SVs in highly repetitive regions) or population-scale screening at the lowest per-genome cost, ONT-only strategies deliver excellent results. Contact our team for a free project consultation and platform recommendation based on your specific research objectives.
| Category | Requirement | Notes |
| Sample type | Genomic DNA (blood, buffy coat, EBV-transformed cell line, fresh or frozen tissue) | DNA extraction service available for challenging sample types (FFPE samples have specific limitations for long-read library prep) |
| Minimum input (gDNA) | 1–5 µg (PacBio HiFi); 1–5 µg (Nanopore) | HMW DNA (≥ 30 kb for PacBio, ≥ 20 kb for Nanopore) strongly recommended for optimal sequencing performance; low-input protocols available for limited samples |
| DNA quality | OD260/280: 1.8–2.0; OD260/230: ≥ 1.8; no visible degradation; HMW confirmed by PFGE or TapeStation | DNA purity is critical for long-read library preparation and sequencing yield; samples with significant contamination may require re-extraction or additional purification |
| Coverage recommendation | 10–15× (SV discovery); 15–25× (comprehensive variant detection); 25–30× (de novo assembly) | Coverage recommendations are adjusted based on specific research objectives; for tumor samples, higher coverage (≥20× tumor + matched normal) is recommended for somatic variant detection |
| Sample numbers | Single sample to cohort-scale studies | Multiplexed barcoding (up to 24-plex per Revio SMRT Cell, up to 96-plex per PromethION flow cell) enables cost-effective sequencing of multiple samples; dedicated project management for large cohorts |
| Shipping conditions | gDNA: ice pack (4°C) or dry ice; Blood/tissue: dry ice | See our Sample Submission Guidelines for detailed instructions on sample preparation and shipping |
Independent Platform Expertise
We operate PacBio HiFi (Revio, Sequel II) and Oxford Nanopore (PromethION R10.4.1) platforms in-house, enabling truly platform-neutral recommendations for every human WGS project. Our team has deep experience across both technologies and understands the specific strengths and trade-offs of each platform for different human genomics applications — from rare disease SV detection to population-scale screening.
Benchmark-Driven Coverage Design
Our coverage recommendations are informed by published benchmark studies — including the Comprehensive in-silico Benchmarking of SV detection (Liu et al. 2024, Nature Communications) and results from the All of Us long-read WGS program — ensuring that each project is designed with coverage targets proven to achieve the required detection sensitivity for its specific variant classes of interest.
Comprehensive Bioinformatics Pipeline
Our analysis pipeline covers the full spectrum of human genome analysis from raw read processing through multi-caller SV detection, tandem repeat analysis, mobile element insertion detection, methylation calling, and pharmacogenomic annotation — all delivered with comprehensive quality metrics and publication-ready reporting.
End-to-End Research Support
Every human WGS project includes dedicated project management, regular progress updates, and comprehensive downstream support including NCBI BioProject submission, methods section writing for publications, and data delivery in standard formats (FASTQ, BAM, VCF, FASTA) compatible with downstream analysis tools and databases.
Fatima N, Petri A, Gyllensten U, Feuk L, Ameur A. Evaluation of Single-Molecule Sequencing Technologies for Structural Variant Detection in Two Swedish Human Genomes. Genes. 2020;11(12):1444.
Structural variants are a major source of genetic diversity in the human genome and are increasingly recognised as contributors to both rare and common diseases. Despite their importance, SV detection has lagged behind SNV/InDel detection — largely because short-read sequencing technologies, which produce reads shorter than most SV breakpoint-spanning distances, detect only a fraction of SVs present in any given genome. Long-read single-molecule sequencing technologies from Pacific Biosciences and Oxford Nanopore Technologies promised to overcome this limitation, but at the time of this study (2020), direct head-to-head comparisons of the two platforms for human WGS SV detection in the same individuals were limited.
In this study, Fatima et al. sequenced two Swedish human genomes (Swe1, male; Swe2, female) on both ONT PromethION (~32× coverage) and PacBio Sequel II (HiFi and CLR modes, ~66× CLR coverage, downsampled to ~33× for comparison) and systematically evaluated the SV detection performance of each platform.
DNA from two Swedish individuals (Swe1 and Swe2, selected from the SweGen project) was sequenced on ONT PromethION (six flow cells per individual, ~32× average coverage) and PacBio Sequel II (eight SMRT Cells per individual, ~66× CLR coverage). PacBio data was also processed in HiFi CCS mode. SV calling was performed using minimap2 for alignment and Sniffles for SV detection, with additional validation using PBSV (PacBio) and NanoSV (ONT). Overlap between platforms was assessed at multiple coverage levels, including downsampling of PacBio data to match ONT coverage (~33×) for fair comparison.
Figure 2. Comparison of structural variant detection by PacBio and ONT in two Swedish human genomes. ONT detected ~17,000 SVs per genome, while PacBio detected ~23,000 SVs per genome at full coverage. Downsampling analysis showed that the majority of platform-specific SVs were detected at lower coverage by the other platform, and that combining both technologies provides the most comprehensive SV discovery. Adapted from Fatima et al. (2020), Genes, CC BY 4.0.
This study demonstrated that ONT and PacBio long-read sequencing have similar overall performance for SV detection in human WGS data, with platform-specific differences driven primarily by read length (ONT advantage for spanning long repeats) and per-base accuracy (PacBio advantage for precise breakpoint resolution). The findings established that coverage depth is the single most important determinant of SV detection sensitivity for both platforms and that combining data from both technologies provides the most comprehensive view of human structural variation. These results continue to inform platform selection and coverage design for human WGS projects today.
CD Genomics provides free project consultation to help determine the optimal human WGS strategy for your specific research questions. Contact our scientists to discuss your project requirements.
Coverage recommendations depend on the research objectives. For population-scale SV discovery and screening, 8–15× PacBio HiFi is sufficient, as demonstrated by large-scale programs such as All of Us. For comprehensive variant detection including SNVs, InDels, and SVs with high sensitivity, 15–25× HiFi or ONT is recommended. For de novo genome assembly and the highest quality variant benchmarking, 25–30× HiFi provides the data needed for haplotype-resolved assembly and nucleotide-resolution variant characterisation. These recommendations are derived from published benchmark studies — we provide specific guidance during project design based on your research questions.
Long-read WGS detects multiple variant classes that short-read sequencing systematically misses or mischaracterises. These include: (1) large insertions (the most common SV type in the human genome, comprising ~50% of all SVs) — short reads cannot span insertion breakpoints; (2) tandem repeat expansions — short reads are typically shorter than the expanded repeat array; (3) complex SVs involving multiple breakpoints (inversions with flanking deletions, templated insertion sequences, chromothripsis); (4) SVs in segmental duplications and other highly repetitive regions where short reads cannot be uniquely mapped; (5) mobile element insertions (Alu, LINE1, SVA) at nucleotide resolution; and (6) structural variation in pharmacogene loci (CYP2D6, etc.) where gene duplication/deletion and hybrid gene formation are common.
Yes. Both PacBio HiFi and Oxford Nanopore sequencing provide genome-wide DNA methylation data from the same sequencing run used for variant detection — without bisulfite conversion or antibody enrichment. PacBio HiFi detects base modifications through differences in polymerase kinetics during SMRT sequencing, identifying 5mC and 6mA. Oxford Nanopore detects modified bases directly through changes in the electrical current signal as DNA passes through the nanopore, enabling detection of 5mC, 5hmC, 6mA, and 4mC. This simultaneous genome and epigenome analysis from a single long-read WGS experiment provides integrated genetic and epigenetic characterisation with no additional sample preparation or sequencing cost.
All human genetic data is handled in accordance with applicable data protection regulations and institutional guidelines. Our data management practices include: secure data transfer protocols, encrypted storage on access-controlled servers, sample anonymization where required by the project, and data retention/deletion policies agreed with the client at project initiation. Data is delivered directly to the client — we do not retain human genome sequence data beyond the agreed project period. Specific data handling requirements can be discussed during project design to ensure compliance with institutional ethics and data governance policies.
Typical turnaround times depend on coverage depth, platform choice, and the scope of bioinformatics analysis. For a standard human WGS project at 15–20× coverage on a single platform with basic bioinformatics analysis: approximately 30–45 working days from sample receipt to data delivery. Projects requiring multi-platform sequencing, deep coverage (≥25×), or advanced bioinformatics (de novo assembly, methylation analysis, pharmacogenomic annotation) may require 45–60 working days. Expedited timelines may be available for urgent projects. A detailed project timeline with milestone dates is provided during the project design phase.
Deliverable Examples for Human Whole-Genome Sequencing Projects
1. Raw sequencing data files (FASTQ, POD5 or BAM) with basecalling quality metrics, coverage statistics, and alignment reports (BAM/CRAM files aligned to GRCh38 or T2T-CHM13 reference).
2. Variant call files (VCF/BCF) for all detected variant classes: SVs (with breakpoint coordinates, SV type, and allele frequency estimates), SNVs/InDels, tandem repeat genotypes, mobile element insertions, and copy number variants — with variant quality scores and annotation.
3. Bioinformatics analysis report including: comprehensive variant summary with classification by type and size, SV overlap analysis between detection methods, genome-wide variant density plots, methylation frequency tracks (if requested), and comparison with known population frequency databases.
4. Publication-ready figures and NCBI submission support — including IGV validation screenshots for key variants, genome-wide SV distribution plots, and BioProject submission files for data release.
Figure 3. Representative deliverable formats for human WGS projects. Left: genome-wide SV distribution and classification. Center: multi-caller SV comparison and validation. Right: comprehensive variant annotation and analysis report. AI-generated representative data.
References
For research use only. Not for use in diagnostic procedures.