Population-Scale HiFi Long-Read Multi-Omics Analysis Service — PacBio HiFi Whole-Genome Sequencing for Cohort-Based Association Studies

Population-Scale HiFi Long-Read Multi-Omics Analysis Service — PacBio HiFi Whole-Genome Sequencing for Cohort-Based Association Studies

CD Genomics offers a dedicated Population-Scale HiFi Long-Read Multi-Omics Analysis Service that combines the throughput power of the PacBio Revio platform with a comprehensive bioinformatics pipeline purpose-built for cohort-based genomic discovery. From a single HiFi whole-genome sequencing dataset at 10–30× coverage, we deliver integrated variant call sets spanning single-nucleotide variants (SNVs), structural variants (SVs), tandem repeats (TRs), copy-number variants (CNVs), haplotype-resolved assemblies, and DNA methylation profiles — together with multi-omics association analysis (TR-xQTL, eQTL, methylation QTL) that connects genetic variation to functional and regulatory mechanisms.

What Makes This Service Different

For Research Use Only. Not for use in diagnostic procedures, clinical decision-making, personal health assessment, or therapeutic decision-making.

Why HiFi Sequencing for Population Cohorts?

Population-scale genomic studies have long relied on short-read sequencing and genotyping arrays, which are effective for SNV discovery but miss the classes of variation most associated with complex disease and regulatory function. Structural variants, tandem repeat expansions, and haplotype-resolved regulatory elements remain largely invisible to short-read methods, creating a systematic blind spot in cohort-based association studies.

PacBio HiFi sequencing addresses this gap with >99.9% single-read accuracy and average read lengths exceeding 15 kb, enabling direct observation of SVs and TRs at nucleotide resolution. A growing body of evidence from large-scale initiatives — including the All of Us Research Program, the Human Pangenome Reference Consortium, and the CoLoRS database — has demonstrated that HiFi sequencing at population scale (>1,000 samples) is technically and economically feasible, and that it uncovers disease-associated variation that short-read sequencing systematically misses. A recent analysis of 1,027 HiFi genomes from the All of Us cohort identified 291 SV-disease associations spanning 226 conditions, with 50.9% of these associations involving SVs that were absent from matched short-read data (All of Us Research Program, medRxiv 2025).

Our service translates these advances into a standardised, scalable offering for researchers who need to move beyond SNV-centric genotyping and capture the full spectrum of genetic variation in their cohort studies.

Applications — SV, TR, CNV, Haplotype, Methylation and Multi-Omics Association

Structural Variant & Tandem Repeat Discovery

  • Comprehensive SV calling (deletions, duplications, inversions, translocations, mobile element insertions) at population scale using HiFi-optimised callers (pbsv, Sniffles2, SVDSS).
  • Tandem repeat genotyping across 1.7+ million loci using TRGT and ExpansionHunter, with repeat length and motif-purity resolution at single-nucleotide level.
  • Population frequency cataloguing of SVs and TRs for cohort-wide association analyses.

Haplotype-Resolved & Multi-Omics Association

  • Haplotype-resolved SNV, SV, and TR calling with read-backed phasing for allele-specific expression (ASE) and regulatory analysis.
  • Integrated eQTL, TWAS, and TR-xQTL analysis (TR-eQTL, TR-sQTL, TR-3'aQTL, TR-edQTL) linking population variation to transcriptomic and epitranscriptomic phenotypes.
  • DNA methylation profiling from HiFi kinetic signals for population-scale EWAS, meQTL, and eQTM analysis.
  • TR-caQTL for chromatin accessibility and TR-hQTL for histone modification association.

Copy-Number Variation & Complex Regions

  • CNV detection in segmental duplications, multigene families, and clinically relevant loci (e.g., CYP2D6, C4, SMN1/2, HLA).
  • Long-read resolved copy-number analysis in regions refractory to short-read mapping.

Pangenome & Evolutionary Genomics

  • Pangenome graph construction and alignment for unbiased variant discovery across diverse populations.
  • Population-specific allele frequency estimation and selection signature analysis.
  • Comparative and evolutionary genomics using haplotype-resolved assemblies.

Available Analysis Modules

Our analysis pipeline is organised into tiered modules. The Basic module covers core variant discovery and genotyping; the Advanced module adds multi-omics integration, pangenome analysis, and custom statistical modelling.

Analysis Module Basic Advanced
HiFi WGS alignment (pbmm2, minimap2)
SNV/indel calling (DeepVariant, GATK)
Structural variant calling (pbsv, Sniffles2, SVDSS)
Tandem repeat genotyping (TRGT, ExpansionHunter HiFi)
CNV detection (HiFiCNV, depth-based)
Haplotype phasing (WhatsHap, HiFi-based)
DNA methylation profiling (5mC from HiFi kinetic signals)
Population variant frequency & QC (HWE, MAF, PCA, IBD)
GWAS / SNP-based association testing (SAIGE, REGENIE)
TR-based GWAS — tandem repeat variant association testing
TR-xQTL repeat-phenotype association
TR-eQTL — TR → gene expression association
TR-meQTL — TR → DNA methylation association
TR-sQTL — TR → mRNA splicing association
TR-3'aQTL — TR → alternative polyadenylation (APA) site usage
TR-caQTL — TR → chromatin accessibility association
TR-hQTL — TR → histone modification association
TR-edQTL — TR → RNA editing association
TWAS — transcriptome-wide association study
EWAS — epigenome-wide association study (DNA methylation)
meQTL — SNP → DNA methylation QTL
eQTM — expression quantitative trait methylation
eQTL integration — SNP → gene expression (GTEx reference)
ASE analysis — allele-specific expression
Pangenome graph analysis (minigraph, vg)
Multi-platform / multi-omics integration (RNA-seq, proteomics, metabolomics)
Transgenerational epigenetic inheritance analysis
Post-GWAS functional interpretation (GWAS + eQTL + Hi-C/3D genome integration)
Machine learning-based variant prioritisation

Recommended Project Packages

We offer three pre-configured project packages designed to match different research goals and cohort sizes. Each package includes HiFi sequencing on the PacBio Revio platform, standard QC, and the corresponding analysis modules.

Feature Population Discovery Multi-Omics
Cohort size 1,000–10,000+ samples 500–2,000 samples 100–500 samples
Coverage 8–12× HiFi 12–18× HiFi 20–30× HiFi
Analysis tier Basic Basic + Advanced Basic + Advanced + custom
Variant scope SNV + SV + TR + CNV + methylation + haplotype + pangenome + multi-omics
Association scope GWAS (SNP-based) GWAS + TR-based GWAS + TR-xQTL + meQTL GWAS + TWAS + EWAS + full TR-xQTL suite (8 TR-xQTL types) + eQTM + ASE
Reference benchmark All of Us HiFi (1,027 @ 8×) CoLoRS database (~1,000 HiFi) HPRC + GIAB Platinum
Deliverables VCF + BAM + QC report + GWAS summary + methylation tracks + haplotype blocks + TR genotypes + TR-GWAS results + pangenome graph + multi-omics integration report + all TR-xQTL results + custom visualisations

All packages include project consultation, sample QC, sequencing, standard bioinformatics, and a final project report. Custom coverage depths, cohort-specific analysis designs, and multi-platform integration (e.g., combining HiFi with ONT ultra-long or RNA-seq) are available on request. Coverage recommendations are informed by published large-cohort HiFi studies: All of Us (8× HiFi, 1,027 samples), CoLoRS (~1,000 HiFi genomes), and HPRC (30× HiFi, ~100 samples).

Workflow — From Sample to Population-Level Biological Insight

1. Sample Preparation & QC

High-molecular-weight (HMW) DNA extraction with quality assessment (pulsed-field gel electrophoresis, Quantus fluorometer, TapeStation). Minimum DNA integrity is verified before library construction.

2. HiFi Library Construction & Sequencing

SMRTbell library preparation with PacBio HiFi chemistry. Sequencing on the PacBio Revio system with SMRT Cell 25M throughput. Real-time circular consensus sequencing (CCS) generates HiFi reads at >Q30 accuracy.

3. Primary Data Processing & QC

Adapter trimming, read quality filtering, HiFi read polishing. Per-sample alignment QC (coverage depth, mapping rate, insert-size distribution). Sample identity and ancestry QC checks (PCA against reference panels).

4. Multi-Variant Calling & Genotyping

Parallel calling of SNVs, SVs, TRs, CNVs, and methylation haplotypes using HiFi-optimised tools. Variant-level QC (GQ, DP, Mendelian concordance for family samples). Population-level QC (HWE, MAF filtering, IBD verification).

5. Multi-Layer Association Analysis & Multi-Omics Integration

Cohort-wide GWAS / pheWAS using SAIGE or REGENIE for each variant class (SNP, SV, TR, CNV). TR-based GWAS and full TR-xQTL suite (TR-eQTL, TR-meQTL, TR-sQTL, TR-3'aQTL, TR-caQTL, TR-hQTL, TR-edQTL) connecting repeat genotypes to molecular phenotypes across transcriptomic, epigenetic, and regulatory layers. TWAS (transcriptome-wide) and EWAS (epigenome-wide) association studies. meQTL, eQTM, and ASE analysis for integrated multi-omics interpretation. Multi-platform integration with RNA-seq, proteomics, or metabolomics data where available.

6. Interpretation & Delivery

Annotated variant call sets, association summary statistics, functional enrichment analysis, and publication-ready visualisations (Manhattan plots, QQ plots, locus zoom plots, TR allele distributions).

End-to-end Population-Scale HiFi Multi-Omics Analysis workflow: from HMW DNA extraction and PacBio Revio HiFi sequencing through multi-variant calling, association analysis, and integrated interpretation.

Deliverables

Deliverable Format Description
Raw HiFi sequencing data BAM (CCS) Demultiplexed, adapter-trimmed HiFi reads
Aligned reads BAM (indexed) Reads aligned to reference genome (GRCh38 / T2T-CHM13)
Variant call sets VCF + GVCF SNV, SV, TR genotypes, and CNV calls
Methylation profiles BED / bigWig Per-sample methylation frequencies (5mC)
Haplotype phasing VCF (phased) + HAP Read-backed phased variant blocks
Association results Summary stats + plots GWAS, TR-based GWAS, TWAS, EWAS, TR-xQTL (8 types), eQTL, meQTL, eQTM results with Manhattan and QQ plots
QC report HTML + PDF Per-sample and cohort-level sequencing and variant QC metrics
Project report PDF Methods, analysis parameters, results summary, and data interpretation

Sample and Data Requirements

Category Requirement Notes
Sample type Blood, buffy coat, frozen tissue, or high-quality gDNA For FFPE or degraded samples, please consult our technical team
Minimum input 5 µg HMW gDNA (recommended 10–20 µg) High-molecular-weight DNA (≥30 kb modal length) strongly recommended
DNA quality OD260/280 1.8–2.0; OD260/230 ≥ 2.0; no visible degradation Assessed by pulsed-field gel electrophoresis and TapeStation
Coverage 10–30× HiFi (depending on package) Higher coverage recommended for SV/TR discovery and haplotype phasing
Data transfer Hard drive, FTP, or cloud transfer Secure transfer arrangements made per project

For projects involving re-sequencing of previously sequenced samples, we accept pre-extracted gDNA or sequencing-ready libraries. Sample shipment guidelines are available on our Sample Submission page.

Case Study — Tandem Repeat Allele Profiling in 1,027 HiFi Genomes Reveals Genome-Wide Patterns of Pathogenicity

Danzi MC, Xu IRL, Fazal S, Dolzhenko E, Pellerin D, Weisburd B, et al. (All of Us Research Program Long Read Working Group). bioRxiv 2025. DOI: 10.1101/2025.01.06.631535. (CC BY 4.0)

1. Background

Tandem repeats (TRs) constitute some of the most polymorphic regions of the human genome and are implicated in over 60 known repeat expansion disorders. However, the full spectrum of human TR variation — including the allele sequences, length distributions, and population frequencies — has remained poorly characterised because short-read sequencing cannot resolve repetitive regions at nucleotide resolution. The All of Us Research Program generated PacBio HiFi whole-genome sequencing data from 1,027 participants, creating an opportunity to profile TR variation at an unprecedented scale.

2. Methods

The study used PacBio HiFi sequencing at ~8× coverage across 1,027 genomes from the All of Us cohort. TR genotyping was performed using the Tandem Repeat Genotyping Tool (TRGT), which leverages HiFi read accuracy to determine repeat length, motif sequence, and motif purity at each of 1.7+ million TR loci genome-wide. The resulting database comprises over 3.6 billion TR allele sequences, making it the largest publicly available HiFi-based TR resource to date.

3. Results

Figure adapted from Danzi et al. 2025 (CC BY 4.0). Tandem repeat allele profiling across 1,027 HiFi genomes reveals population-level repeat purity, length distributions, and known pathogenic repeat loci as the most variable loci genome-wide.

Key Findings

  • TR constraint metric. A novel "tandem repeat constraint" measure was developed to distinguish potentially pathogenic from benign TR loci, analogous to the pLI metric used for protein-coding genes.
  • Novel pathogenic expansions. Two novel candidate pathogenic repeat expansions were identified and validated, demonstrating the discovery power of population-scale HiFi TR genotyping.
  • Known pathogenic loci are the most variable. Known repeat expansion loci (e.g., HTT, C9orf72, DMPK, FMR1) ranked among the most variable TRs genome-wide when measured by longest pure motif segment length, validating the approach.
  • Population-specific patterns. Extensive population-specific TR allele distributions were observed across ancestral groups represented in the All of Us cohort.

4. Conclusions

This study demonstrates that population-scale HiFi TR genotyping is technically robust and biologically informative. The 3.6-billion-allele TR resource provides a critical reference for future repeat-disease association studies and establishes a framework for integrating TR variation into population genomics pipelines. It underscores the value of HiFi sequencing as a platform for comprehensive variant discovery in large cohorts.

Frequently Asked Questions

Demo & Talk to a Long-Read Population Genomics Specialist

Our demo package includes representative deliverables from a pilot cohort HiFi WGS project:

1. Multi-variant call set browser view — SNV, SV, TR, CNV, and methylation tracks for a single sample displayed in IGV.

2. Cohort-level association summary — Manhattan plot, QQ plot, and TR allele distribution across the pilot cohort.

3. TR-xQTL results — example TR-phenotype association with locus zoom plot and repeat motif visualisation.

4. TWAS and EWAS results — transcriptome-wide and epigenome-wide association signals with gene-set enrichment and pathway analysis.

5. Multi-omics integration network — circos plot showing interactions between TRs, SNPs, gene expression, DNA methylation, and phenotype across QTL types.

Representative deliverable format: multi-variant association summary, TR allele distribution, and population-level QC metrics from a cohort HiFi WGS project.

Population-scale long-read projects require careful experimental design. Our scientific team — experienced in large-cohort HiFi sequencing, multi-variant calling, and multi-omics integration — can help you scope the optimal coverage, analysis modules, and project timeline for your specific research goals.

Submit an inquiry → or contact us at contact@cd-genomics.com | (631) 259-7705

References

  1. Danzi MC, Xu IRL, Fazal S, et al. Detailed tandem repeat allele profiling in 1,027 long-read genomes reveals genome-wide patterns of pathogenicity. bioRxiv. 2025. DOI: 10.1101/2025.01.06.631535. (CC BY 4.0)
  2. Wen H, Yang J, Zhao X, et al. TRFill: synergistic use of HiFi and Hi-C sequencing enables accurate assembly of tandem repeats for population-level analysis. Genome Biology. 2025;26:227. DOI: 10.1186/s13059-025-03685-5. (CC BY 4.0)
  3. Weisburd B, Dolzhenko E, Bennett MF, et al. Defining a tandem repeat catalog and variation clusters for genome-wide analyses and population databases. Am J Hum Genet. 2026. DOI: 10.1016/j.ajhg.2026.03.020.
  4. All of Us Research Program Long Read Working Group. Population-scale Long-read Sequencing in the All of Us Research Program. medRxiv. 2025. DOI: 10.1101/2025.10.02.25336942. (CC BY 4.0)

Related Services

Get Your Instant Quote