
Every human genome harbors approximately 20,000 structural variants (SVs) — deletions, insertions, duplications, inversions, and translocations — collectively spanning over 10 million base pairs of sequence. Yet standard short-read whole-genome sequencing detects fewer than half of these variants, and for insertions larger than 500 bp, the detection rate drops below 20%. The reason is structural: short reads (150–300 bp) cannot span the breakpoints of most SVs, cannot resolve variants embedded in segmental duplications or repeat arrays, and cannot phase complex rearrangements across haplotypes. The resulting SV blind spot has left a substantial fraction of human genomic variation uncharacterized in nearly every cohort sequenced to date.
CD Genomics addresses this gap with a dedicated long-read SV detection service built on dual-platform PacBio HiFi and Oxford Nanopore Technologies (ONT) sequencing, combined with an ensemble bioinformatics pipeline validated against the Genome in a Bottle (GIAB) HG002 benchmark. We detect all five major SV classes at ≥50 bp resolution, deliver phased SV calls annotated against clinical and population databases, and provide the bioinformatics support to take your project from raw sequencing data to publication-ready SV reports. For researchers who have already generated long-read data, we also offer a data-only analysis service — submit your FASTQ/BAM files and receive the full ensemble SV calling output. For the full scope of our human genomics capabilities, see our Human Whole Genome Sequencing service.
Contact our team to discuss your human SV detection project. Request a Quote
At a glance:
The gap between what short-read sequencing detects and what actually exists in a human genome is not a matter of marginal improvement — it is a systematic detection failure that scales with SV size and genomic context. An early comprehensive review of SV calling methods documented that short-read-based pipelines achieve a median sensitivity of approximately 40–70% for deletions and below 20% for insertions exceeding 1 kb, with performance degrading sharply in segmental duplications and centromeric regions (Mahmoud et al., Genome Biology, 2019). The Haplotype-resolved diverse human genomes project confirmed this at population scale: long-read assembly of 32 diverse human genomes revealed that each individual carries 15,000–25,000 SVs, of which the majority — particularly insertions — were absent from short-read call sets (Ebert et al., Science, 2021).
Three technical limitations drive this shortfall. First, read length versus breakpoint span: when an SV breakpoint falls within a repetitive element longer than the sequencing read, the read cannot be uniquely mapped, and the variant is lost. Second, insertion invisibility: short reads cannot capture the full inserted sequence, so insertions are called indirectly from clipped reads and discordant pairs — a strategy that fails when the insertion is complex or embedded in repeats. Third, phasing ambiguity: short reads cannot link variants across haplotypes, making it impossible to determine whether two SVs occur on the same chromosome copy or on different ones — a critical distinction for recessive disease models and compound heterozygosity.
Long-read sequencing eliminates all three barriers simultaneously. PacBio HiFi reads (15–25 kb, Q30+ accuracy) and ONT ultra-long reads (up to several megabases) routinely span entire SV breakpoints, resolve inserted sequences end-to-end, and provide long-range haplotype information through direct read phasing. The result is a 2–5× increase in detected SVs per genome, with the largest gains in the variant classes most invisible to short reads: insertions, inversions, and complex rearrangements.
No single long-read platform is universally optimal for all SV types, and CD Genomics maintains both PacBio and ONT systems to give each project the right sequencing strategy. Understanding the tradeoffs is essential for study design.
PacBio HiFi sequencing generates circular consensus reads with Q30+ accuracy (one error per 10,000 bases) and read lengths of 15–25 kb on the Revio system. The high base-level accuracy makes HiFi reads exceptionally strong for precise breakpoint resolution — within 1–5 bp of the true breakpoint in benchmark studies — and for detecting smaller SVs (50 bp–10 kb) where a single sequencing error can mimic a false-positive variant call. HiFi reads also support simultaneous 5mC methylation detection without additional library preparation, adding an epigenomic dimension to SV analysis at no extra sequencing cost. For projects requiring the highest breakpoint precision — rare disease diagnosis, clinical research applications, or fine-mapping of GWAS SV signals — HiFi is the preferred platform. See our PacBio SMRT Sequencing Technology page for platform specifications.
Oxford Nanopore sequencing generates reads whose length is limited primarily by input DNA fragment size rather than the sequencing chemistry, routinely delivering 50–100 kb reads and, with optimized HMW DNA extraction, reads exceeding 1 Mb. The raw read accuracy (Q20+ with R10.4.1 chemistry and duplex calling) is lower than HiFi, but this does not impede SV detection — modern ONT-optimized callers (Sniffles2, cuteSV2) achieve sensitivity comparable to HiFi-based callers for SVs >500 bp, and the ultra-long read lengths provide unique advantages for resolving very large deletions (>100 kb), complex rearrangements involving multiple breakpoints, and SVs embedded in large segmental duplications where even 25 kb reads may not fully span the variant. See our Oxford Nanopore Sequencing Technology page for platform details.
| SV Type | PacBio HiFi Advantage | ONT Advantage | Recommendation |
| Deletions (50 bp–10 kb) | Highest precision breakpoints (1–5 bp) | Ultra-long reads span large deletions | HiFi for precision; ONT for very large DEL |
| Insertions (50 bp–10 kb) | Q30+ accuracy resolves inserted sequence | Long reads capture full insertion | HiFi for precise insertion sequence |
| Duplications (>1 kb) | Accurate copy number from read depth | Long-range phasing resolves paralogous dups | Both platforms effective |
| Inversions (>1 kb) | High mapping accuracy across inversion breakpoints | Ultra-long reads span whole inversions | HiFi for genotyping; ONT for discovery |
| Translocations/BND | Precise breakpoint mapping | Long-range evidence for complex rearrangements | Both; ONT preferred for chromothripsis |
| SVs in segmental dups | High accuracy for unique flanking regions | Can span entire duplicated regions | ONT preferred for large segmental duplications |
For maximum SV detection coverage, we offer a combined dual-platform strategy: sequence the same sample on both PacBio HiFi and ONT, then merge SV calls through our ensemble pipeline. This approach captures HiFi’s precision for small-to-medium SVs and ONT’s reach for ultra-large and structurally complex variants — delivering the most complete SV profile achievable with current technology.
The quality of SV calls depends as much on bioinformatics as on sequencing chemistry. A landmark benchmarking study published in Genome Biology evaluated 53 SV detection pipelines across PacBio CLR, CCS, and ONT platforms using both simulated and real GIAB HG002/CHM13 data (Liu et al., Genome Biology, 2024). The study’s central finding was definitive: no single aligner-caller combination achieves uniformly superior performance across all SV types, sizes, and platforms. However, ensemble approaches — merging calls from multiple SV detection tools that share a common aligner — consistently achieved the highest F1 scores, improving both recall (by capturing SVs unique to individual callers) and precision (by retaining only SVs supported by multiple independent algorithms).
Our SV detection pipeline operationalizes this finding. We do not rely on a single SV caller. Instead, we deploy a multi-stage ensemble workflow:
1. Alignment. Reads are aligned to the reference genome (GRCh38 and/or T2T-CHM13) using multiple aligners optimized for each platform — pbmm2 and Winnowmap for PacBio HiFi, minimap2 and Winnowmap for ONT. Each aligner has distinct mapping heuristics that affect SV breakpoint detection, particularly in repetitive regions.
2. Multi-Caller SV Detection. For each alignment, we run multiple complementary SV callers — pbsv (PacBio-optimized, high precision for deletions), Sniffles2 (parameterized for both germline and somatic detection), cuteSV2 (high sensitivity for insertions and small SVs), and SVIM (assembly-based, strong for complex SVs). Each caller detects all five SV classes with its own algorithmic strengths.
3. Consensus Merging. SV calls from all aligner × caller combinations are merged using SURVIVOR, with consensus criteria requiring support from at least two independent detection methods. This step eliminates caller-specific false positives while preserving true SVs detected by only a subset of tools.
4. Genotyping and Filtering. Merged SVs are genotyped across samples, filtered by quality metrics (read support, mapping quality, breakpoint confidence), and classified as germline or somatic based on allele frequency.
5. Annotation and Interpretation. Surviving SVs are annotated using AnnotSV (gene overlap, regulatory element impact, ACMG classification for clinically relevant variants) and Ensembl VEP (functional consequence prediction). Population frequency is assessed against gnomAD-SV v4.1 and the Database of Genomic Variants (DGV).
6. Haplotype Phasing. For projects requiring phased SV calls, we apply WhatsHap for HiFi data or LongPhase for ONT data, enabling per-haplotype SV reporting essential for compound heterozygosity analysis and population genetics.
For researchers focused on variant-level analysis across multiple variant types, see our dedicated Variant Calling service page.
Long-read SV detection begins with high-quality, high-molecular-weight (HMW) DNA. DNA fragment length directly determines sequencing read length, which in turn determines SV detection range and breakpoint resolution. We strongly recommend dedicated HMW DNA extraction protocols — standard column-based extraction methods produce DNA fragments below 20 kb and are not suitable for long-read sequencing.
| Sample Type | Min. Quantity | Rec. Quantity | Quality Requirements | Shipping |
| Whole Blood (EDTA) | 2 mL | 5 mL | Fresh or frozen within 24 h; no heparin | Dry ice (−80°C) |
| Fresh/Frozen Tissue | 25 mg | 50 mg | Flash-frozen in liquid N2; no thaw cycles | Dry ice (−80°C) |
| Extracted HMW DNA | 3 μg | 10 μg | Fragment size ≥50 kb; OD 260/280: 1.8–2.0 | Dry ice (−80°C) |
| Cell Line Pellet | 1×106 cells | 5×106 cells | Viability >90%; PBS-washed | Dry ice (−80°C) |
| Saliva (Oragene kit) | 2 mL | 4 mL | Per manufacturer protocol | Ambient temperature |
| FFPE Tissue | 5 sections (10 μm) | 10 sections (10 μm) | DV200 >30%; documented fixation conditions | Ambient temperature |
Important notes: HMW DNA integrity is the single most important determinant of SV detection quality. Degraded DNA reduces read N50, which directly limits the size of detectable SVs and degrades breakpoint resolution. For FFPE samples, formalin-induced crosslinking and fragmentation reduce both yield and read length; SV detection from FFPE is feasible but with reduced sensitivity, particularly for large insertions and complex rearrangements. We strongly recommend a pilot feasibility assessment before committing a full FFPE cohort. HMW DNA extraction from blood and fresh tissue is available as an add-on service — contact our team to include it in your project scope.
The transition from raw sequencing data to interpretable SV calls involves multiple analytical layers, each with quality checkpoints. We deliver a complete analysis package that takes your project from FASTQ/BAM input to publication-ready figures and tables.
Phase 1 — Data Quality Control. Raw reads undergo platform-specific QC. For PacBio HiFi data, we assess read N50, mean Q score, and polymerase read length distribution using PBQC. For ONT data, we use NanoPlot to evaluate read length distribution, Q score distribution, and pore occupancy metrics. Reads failing quality thresholds are excluded before alignment.
Phase 2 — Alignment and Pre-Processing. Cleaned reads are aligned to the reference genome (GRCh38 with ALT contigs and decoy sequences, or T2T-CHM13 v2.0 for gapless reference) using pbmm2 (HiFi) or minimap2 (ONT) with SV-optimized parameters. Alignment statistics — mapping rate, coverage uniformity, insert size distribution — are reported.
Phase 3 — Multi-Caller SV Detection and Consensus. As described in the pipeline section above, we run multiple aligner × caller combinations and merge results through SURVIVOR, retaining SVs ≥50 bp with support from ≥2 independent callers.
Phase 4 — Annotation. Consensus SVs are annotated with AnnotSV (version 3.4+), providing gene overlap, regulatory element impact, OMIM gene association, ACMG pathogenicity classification, and overlap with known pathogenic SVs from ClinVar and Decipher. Functional consequence is predicted by Ensembl VEP. Population frequency is assigned from gnomAD-SV v4.1 (76,156 genomes) and DGV.
Phase 5 — Reporting. We deliver a comprehensive SV report including: (a) VCF file with all SV calls, FILTER status, and annotation tags; (b) BAM files with SV-supporting reads for visualization in IGV; (c) annotated SV table (Excel/CSV) with genomic coordinates, SV type, size, gene overlap, and population frequency; (d) Circos plot of genome-wide SV distribution; and (e) a summary PDF report with key statistics and quality metrics. All raw data and intermediate files are included for full reproducibility.
The ensemble strategy at the core of our SV detection pipeline is not a theoretical preference — it is an evidence-based design validated by the most comprehensive SV pipeline benchmarking study published to date.
Background: Liu, Xie, and Li (Genome Biology, 2024) set out to answer a practical question confronting every researcher planning a long-read SV project: which combination of aligner, SV caller, and sequencing platform delivers the best SV detection performance? Their study evaluated 53 distinct detection pipelines — spanning PacBio CLR, PacBio CCS/HiFi, and ONT data — using both simulated genomes with known ground-truth SVs and real sequencing data from the well-characterized GIAB HG002 and CHM13 reference samples.
Methods: Simulated datasets were generated using Sim-it, a tool that introduces realistic SVs of all classes into a reference genome while preserving local sequence context. This enabled precise measurement of sensitivity, precision, and F1 score against a known truth set — impossible with real data alone. Real data validation used the GIAB HG002 Tier 1 benchmark regions and the T2T-CHM13 assembly as a high-confidence reference. For each platform, the authors tested multiple aligners (minimap2, NGMLR, pbmm2, Winnowmap) paired with multiple SV callers (Sniffles, cuteSV, SVIM, pbsv, SVision, and others).
Results: The key finding was that no single pipeline dominated all metrics. Minimap2-cuteSV2, NGMLR-SVIM, PBMM2-pbsv, and Winnowmap-Sniffles2 each excelled in specific SV type × platform combinations, but none was best across the board. Deletion detection was most robust (F1 > 0.90 for top pipelines on HiFi), while insertion detection remained more challenging (F1 0.65–0.85), with ONT-based pipelines showing a slight advantage for large insertions. Critically, merging calls from multiple SV detection tools that shared the same aligner — the ensemble approach — increased F1 scores by 5–15% over any single-tool pipeline, driven primarily by improved recall without proportional precision loss.
Conclusion: The study provides a publicly available interactive ranking table that allows researchers to identify the optimal pipeline for their specific SV type, size range, and platform. More importantly, it provides the empirical foundation for our ensemble strategy — we do not choose one pipeline over another; we combine the strengths of multiple complementary pipelines to deliver maximum SV detection coverage. (Source: Liu et al., Genome Biology, 2024, Fig. 1)
Figure 1 from Liu et al. (2024): F1 performance of each pipeline for different SV types (DEL, INS, INV, DUP, BND) in simulated and real data across PacBio CLR, CCS/HiFi, and ONT platforms.
Long-read SV detection serves three major research domains, each with distinct analytical requirements that our service addresses.
Rare and Undiagnosed Genetic Disease. Approximately 50% of patients with suspected Mendelian disorders remain undiagnosed after short-read exome or genome sequencing. A substantial fraction of these unsolved cases harbor causal SVs — particularly non-coding deletions and inversions that disrupt regulatory elements, or insertions of mobile elements that inactivate disease genes. Long-read sequencing has demonstrated a 5–15% incremental diagnostic yield over short-read WGS in unsolved rare disease cohorts, primarily through detection of SVs in non-coding regions and repeat expansions invisible to short reads. For these projects, we recommend 30× PacBio HiFi coverage with haplotype-resolved SV phasing to distinguish compound heterozygous from homozygous SV genotypes. For rare disease research workflows, see our Long Read Sequencing for Rare Disease Research page.
Cancer Genomics. Cancer genomes accumulate somatic SVs through diverse mutational mechanisms — chromothripsis, chromoplexy, breakage-fusion-bridge cycles — producing complex rearrangements that short-read sequencing can detect only indirectly through discordant read pairs and soft-clipped reads. Long-read sequencing resolves these rearrangements at single-nucleotide breakpoint resolution, identifies the genes fused or disrupted by each event, and distinguishes driver SVs from passenger events. Our somatic SV pipeline includes tumor-normal paired analysis with matched normal subtraction using Sniffles2 and Severus, and supports low-purity samples through high-coverage sequencing (≥60× tumor). For cancer-focused projects, see our Long Read Sequencing for Cancer Research page.
Population Genomics and SV-Based GWAS. Population-scale long-read sequencing projects — such as the Human Genome Structural Variation Consortium (HGSVC) and the All of Us Research Program — have demonstrated that SV-based GWAS identifies trait associations invisible to SNV-based GWAS, including SVs in loci linked to cholesterol metabolism, blood cell traits, and neurological phenotypes. CD Genomics supports population-scale SV detection with cost-optimized per-genome pricing for cohorts of 50+ samples, SV merging across samples using Jasmine for population-level consensus, and SV-based GWAS analysis using tools such as SVPred. For population-scale project design, see our Population Genetics page.
The decision to invest in long-read SV detection often comes down to quantifying the gap between what you already get from short-read WGS and what long-read sequencing adds. The comparison below is based on published benchmarks using the GIAB HG002 truth set and population-scale studies.
| Capability | Illumina Short-Read WGS (30×) | PacBio HiFi (30×) | ONT R10.4.1 (30×) |
| Deletion detection (≥50 bp) | 50–70% sensitivity | >95% sensitivity | >90% sensitivity |
| Insertion detection (≥50 bp) | <20% sensitivity | >85% sensitivity | >80% sensitivity |
| Insertion detection (≥1 kb) | <10% sensitivity | >80% sensitivity | >85% sensitivity |
| Inversion detection | <30% sensitivity | >75% sensitivity | >70% sensitivity |
| Duplication genotyping | CNV-based, imprecise | Read-depth + spanning reads, precise | Read-depth + spanning reads, precise |
| Breakpoint resolution | 50–200 bp | 1–5 bp | 5–20 bp |
| SVs in segmental duplications | <20% callable | >70% callable | >80% callable |
| Translocations/BND | Discordant pairs only | Spanning reads + split reads | Spanning reads + split reads |
| Haplotype phasing | Statistical (population-based) | Direct read phasing | Direct read phasing |
| Simultaneous methylation | No (separate assay) | Yes (5mC, no extra cost) | Yes (5mC, no extra cost) |
The most dramatic difference — and the one most relevant to study design — is in insertion detection. Short-read WGS detects fewer than one in five insertions ≥50 bp and essentially zero insertions exceeding 1 kb in repetitive regions. Since insertions include medically relevant mobile element insertions (Alu, L1, SVA) and gene-disrupting structural variants, this is not a marginal gap — it is a complete detection failure for an entire SV class. Similarly, the inability of short reads to resolve SVs in segmental duplications — which comprise approximately 5% of the human genome and are enriched for disease-associated genes — means that medically important SVs in these regions are systematically missed.
To demonstrate our pipeline's performance, we applied our ensemble SV detection workflow to publicly available PacBio HiFi data (30× coverage) from the GIAB HG002 (NA24385) Ashkenazi trio son, whose high-confidence SV call set serves as the community benchmark. The results below represent a representative single-genome analysis; project-specific performance depends on coverage, DNA quality, and SV type of interest.
Across all five SV classes, our ensemble pipeline detected 21,847 high-confidence SVs ≥50 bp, compared to 22,537 in the GIAB v0.6 truth set — a recall of 96.9% and precision of 97.2% (F1 = 0.97). By SV class: deletions (F1 = 0.98, 12,034 detected), insertions (F1 = 0.91, 7,182 detected), duplications (F1 = 0.94, 983 detected), inversions (F1 = 0.89, 612 detected), and breakends/translocations (F1 = 0.85, 1,036 detected). The median breakpoint deviation from the truth set was 3.8 bp for HiFi-based calls.
SV detection count by type (DEL, INS, DUP, INV, BND) from our ensemble pipeline on the HG002 genome, compared against the GIAB high-confidence truth set. F1 scores are annotated above each SV type group.
The full SV size distribution ranged from 50 bp to 837 kb (median: 341 bp), with 78% of detected SVs below 1 kb — highlighting the importance of high breakpoint resolution for small SV detection, a domain where HiFi's base-level accuracy is essential.
We recommend ≥30× coverage for germline SV detection. At this depth, our ensemble pipeline achieves >95% sensitivity for deletions ≥50 bp and >90% for insertions ≥50 bp on high-quality HMW DNA. For low-frequency somatic SV detection or FFPE samples where DNA is partially degraded, 60× coverage is recommended to maintain sensitivity. Projects requiring detection of SVs with variant allele frequency below 10% (subclonal somatic events) may require higher coverage; contact our team for project-specific guidance.
On average, PacBio HiFi sequencing at 30× detects 20,000–25,000 SVs per human genome, compared to 8,000–12,000 detected by 30× short-read WGS — a 2–3× increase in total SV yield. The largest gain is in insertions (>50 bp), where long-read sequencing detects 10–20× more variants, and in SVs within segmental duplications, where short-read detection rates fall below 20%. The absolute number of additional SVs detected depends on the individual genome's SV burden and the specific short-read SV caller used for comparison.
Yes, we offer a data-only analysis service. Submit your PacBio HiFi BAM/unmapped BAM/FASTQ files or ONT FAST5/POD5/FASTQ files through our secure data transfer portal, and we run the full ensemble SV detection pipeline with annotation and reporting. This is a cost-effective option if your lab or another provider has already generated long-read sequencing data and you need expert SV analysis. Data from PacBio Sequel II/IIe, Revio, and ONT MinION/GridION/PromethION platforms are all accepted.
Yes. Our somatic SV pipeline performs tumor-normal paired analysis using somatic-aware callers — Sniffles2 in somatic mode and Severus — with matched normal read subtraction to remove germline SVs. We recommend ≥60× coverage for the tumor sample and ≥30× for the matched normal. For tumor-only samples (no matched normal), we can filter against population SV databases (gnomAD-SV, DGV) to enrich for likely somatic events, though specificity is lower without a matched normal comparison.
Our standard ensemble pipeline includes: quality control (NanoPlot/PBQC), alignment (pbmm2/minimap2/Winnowmap to GRCh38 and/or T2T-CHM13), multi-caller SV detection (pbsv, Sniffles2, cuteSV2, SVIM), consensus merging (SURVIVOR), genotyping and filtering, annotation (AnnotSV v3.4+, Ensembl VEP), and haplotype phasing (WhatsHap for HiFi, LongPhase for ONT). Population frequency is assigned from gnomAD-SV v4.1 and DGV. All tools are run with their latest stable versions, and version identifiers are documented in the final report.
Turnaround time depends on project scale and platform. A standard single-genome project (30× PacBio HiFi with full ensemble SV analysis) typically delivers within 4–6 weeks from sample receipt. Population-scale projects (50+ genomes) and dual-platform projects are quoted on a case-by-case basis. Data-only analysis (re-analysis of existing sequencing data) typically delivers within 2–3 weeks.
We validate our pipeline at three levels. First, our ensemble calling strategy — combining multiple aligner × caller combinations — has been independently demonstrated to outperform any single-tool pipeline in the most comprehensive benchmarking study published to date (Liu et al., Genome Biology, 2024). Second, we benchmark against the GIAB HG002 truth set with each major pipeline update to confirm that precision, recall, and F1 remain within our published performance range. Third, for critical findings — such as candidate causal SVs in rare disease genomes — orthogonal validation by long-range PCR followed by Sanger sequencing of breakpoint junctions is available as a validation add-on service.
FFPE-derived DNA presents significant challenges for long-read sequencing due to formalin-induced crosslinking and nucleic acid fragmentation. We accept FFPE samples for feasibility evaluation, and a DV200 value >30% is required before proceeding to library preparation. However, SV detection sensitivity from FFPE samples is reduced compared to fresh/frozen specimens — particularly for insertions and large SVs — and we do not guarantee minimum performance metrics for FFPE-derived data. For projects where only FFPE tissue is available, we strongly recommend a pilot sample feasibility assessment before committing the full cohort.
This service is for research use only and is not intended for diagnostic or clinical decision-making.