CD Genomics offers a dedicated Population-Scale HiFi Long-Read Multi-Omics Analysis Service that combines the throughput power of the PacBio Revio platform with a comprehensive bioinformatics pipeline purpose-built for cohort-based genomic discovery. From a single HiFi whole-genome sequencing dataset at 10–30× coverage, we deliver integrated variant call sets spanning single-nucleotide variants (SNVs), structural variants (SVs), tandem repeats (TRs), copy-number variants (CNVs), haplotype-resolved assemblies, and DNA methylation profiles — together with multi-omics association analysis (TR-xQTL, eQTL, methylation QTL) that connects genetic variation to functional and regulatory mechanisms.
For Research Use Only. Not for use in diagnostic procedures, clinical decision-making, personal health assessment, or therapeutic decision-making.
Population-scale genomic studies have long relied on short-read sequencing and genotyping arrays, which are effective for SNV discovery but miss the classes of variation most associated with complex disease and regulatory function. Structural variants, tandem repeat expansions, and haplotype-resolved regulatory elements remain largely invisible to short-read methods, creating a systematic blind spot in cohort-based association studies.
PacBio HiFi sequencing addresses this gap with >99.9% single-read accuracy and average read lengths exceeding 15 kb, enabling direct observation of SVs and TRs at nucleotide resolution. A growing body of evidence from large-scale initiatives — including the All of Us Research Program, the Human Pangenome Reference Consortium, and the CoLoRS database — has demonstrated that HiFi sequencing at population scale (>1,000 samples) is technically and economically feasible, and that it uncovers disease-associated variation that short-read sequencing systematically misses. A recent analysis of 1,027 HiFi genomes from the All of Us cohort identified 291 SV-disease associations spanning 226 conditions, with 50.9% of these associations involving SVs that were absent from matched short-read data (All of Us Research Program, medRxiv 2025).
Our service translates these advances into a standardised, scalable offering for researchers who need to move beyond SNV-centric genotyping and capture the full spectrum of genetic variation in their cohort studies.
Our analysis pipeline is organised into tiered modules. The Basic module covers core variant discovery and genotyping; the Advanced module adds multi-omics integration, pangenome analysis, and custom statistical modelling.
| Analysis Module | Basic | Advanced |
| HiFi WGS alignment (pbmm2, minimap2) | ✓ | ✓ |
| SNV/indel calling (DeepVariant, GATK) | ✓ | ✓ |
| Structural variant calling (pbsv, Sniffles2, SVDSS) | ✓ | ✓ |
| Tandem repeat genotyping (TRGT, ExpansionHunter HiFi) | ✓ | ✓ |
| CNV detection (HiFiCNV, depth-based) | ✓ | ✓ |
| Haplotype phasing (WhatsHap, HiFi-based) | ✓ | ✓ |
| DNA methylation profiling (5mC from HiFi kinetic signals) | ✓ | ✓ |
| Population variant frequency & QC (HWE, MAF, PCA, IBD) | ✓ | ✓ |
| GWAS / SNP-based association testing (SAIGE, REGENIE) | ✓ | ✓ |
| TR-based GWAS — tandem repeat variant association testing | — | ✓ |
| TR-xQTL repeat-phenotype association | — | ✓ |
| TR-eQTL — TR → gene expression association | — | ✓ |
| TR-meQTL — TR → DNA methylation association | — | ✓ |
| TR-sQTL — TR → mRNA splicing association | — | ✓ |
| TR-3'aQTL — TR → alternative polyadenylation (APA) site usage | — | ✓ |
| TR-caQTL — TR → chromatin accessibility association | — | ✓ |
| TR-hQTL — TR → histone modification association | — | ✓ |
| TR-edQTL — TR → RNA editing association | — | ✓ |
| TWAS — transcriptome-wide association study | — | ✓ |
| EWAS — epigenome-wide association study (DNA methylation) | — | ✓ |
| meQTL — SNP → DNA methylation QTL | — | ✓ |
| eQTM — expression quantitative trait methylation | — | ✓ |
| eQTL integration — SNP → gene expression (GTEx reference) | — | ✓ |
| ASE analysis — allele-specific expression | — | ✓ |
| Pangenome graph analysis (minigraph, vg) | — | ✓ |
| Multi-platform / multi-omics integration (RNA-seq, proteomics, metabolomics) | — | ✓ |
| Transgenerational epigenetic inheritance analysis | — | ✓ |
| Post-GWAS functional interpretation (GWAS + eQTL + Hi-C/3D genome integration) | — | ✓ |
| Machine learning-based variant prioritisation | — | ✓ |
We offer three pre-configured project packages designed to match different research goals and cohort sizes. Each package includes HiFi sequencing on the PacBio Revio platform, standard QC, and the corresponding analysis modules.
| Feature | Population | Discovery | Multi-Omics |
| Cohort size | 1,000–10,000+ samples | 500–2,000 samples | 100–500 samples |
| Coverage | 8–12× HiFi | 12–18× HiFi | 20–30× HiFi |
| Analysis tier | Basic | Basic + Advanced | Basic + Advanced + custom |
| Variant scope | SNV + SV + TR + CNV | + methylation + haplotype | + pangenome + multi-omics |
| Association scope | GWAS (SNP-based) | GWAS + TR-based GWAS + TR-xQTL + meQTL | GWAS + TWAS + EWAS + full TR-xQTL suite (8 TR-xQTL types) + eQTM + ASE |
| Reference benchmark | All of Us HiFi (1,027 @ 8×) | CoLoRS database (~1,000 HiFi) | HPRC + GIAB Platinum |
| Deliverables | VCF + BAM + QC report + GWAS summary | + methylation tracks + haplotype blocks + TR genotypes + TR-GWAS results | + pangenome graph + multi-omics integration report + all TR-xQTL results + custom visualisations |
All packages include project consultation, sample QC, sequencing, standard bioinformatics, and a final project report. Custom coverage depths, cohort-specific analysis designs, and multi-platform integration (e.g., combining HiFi with ONT ultra-long or RNA-seq) are available on request. Coverage recommendations are informed by published large-cohort HiFi studies: All of Us (8× HiFi, 1,027 samples), CoLoRS (~1,000 HiFi genomes), and HPRC (30× HiFi, ~100 samples).
High-molecular-weight (HMW) DNA extraction with quality assessment (pulsed-field gel electrophoresis, Quantus fluorometer, TapeStation). Minimum DNA integrity is verified before library construction.
SMRTbell library preparation with PacBio HiFi chemistry. Sequencing on the PacBio Revio system with SMRT Cell 25M throughput. Real-time circular consensus sequencing (CCS) generates HiFi reads at >Q30 accuracy.
Adapter trimming, read quality filtering, HiFi read polishing. Per-sample alignment QC (coverage depth, mapping rate, insert-size distribution). Sample identity and ancestry QC checks (PCA against reference panels).
Parallel calling of SNVs, SVs, TRs, CNVs, and methylation haplotypes using HiFi-optimised tools. Variant-level QC (GQ, DP, Mendelian concordance for family samples). Population-level QC (HWE, MAF filtering, IBD verification).
Cohort-wide GWAS / pheWAS using SAIGE or REGENIE for each variant class (SNP, SV, TR, CNV). TR-based GWAS and full TR-xQTL suite (TR-eQTL, TR-meQTL, TR-sQTL, TR-3'aQTL, TR-caQTL, TR-hQTL, TR-edQTL) connecting repeat genotypes to molecular phenotypes across transcriptomic, epigenetic, and regulatory layers. TWAS (transcriptome-wide) and EWAS (epigenome-wide) association studies. meQTL, eQTM, and ASE analysis for integrated multi-omics interpretation. Multi-platform integration with RNA-seq, proteomics, or metabolomics data where available.
Annotated variant call sets, association summary statistics, functional enrichment analysis, and publication-ready visualisations (Manhattan plots, QQ plots, locus zoom plots, TR allele distributions).
End-to-end Population-Scale HiFi Multi-Omics Analysis workflow: from HMW DNA extraction and PacBio Revio HiFi sequencing through multi-variant calling, association analysis, and integrated interpretation.
| Deliverable | Format | Description |
| Raw HiFi sequencing data | BAM (CCS) | Demultiplexed, adapter-trimmed HiFi reads |
| Aligned reads | BAM (indexed) | Reads aligned to reference genome (GRCh38 / T2T-CHM13) |
| Variant call sets | VCF + GVCF | SNV, SV, TR genotypes, and CNV calls |
| Methylation profiles | BED / bigWig | Per-sample methylation frequencies (5mC) |
| Haplotype phasing | VCF (phased) + HAP | Read-backed phased variant blocks |
| Association results | Summary stats + plots | GWAS, TR-based GWAS, TWAS, EWAS, TR-xQTL (8 types), eQTL, meQTL, eQTM results with Manhattan and QQ plots |
| QC report | HTML + PDF | Per-sample and cohort-level sequencing and variant QC metrics |
| Project report | Methods, analysis parameters, results summary, and data interpretation |
| Category | Requirement | Notes |
| Sample type | Blood, buffy coat, frozen tissue, or high-quality gDNA | For FFPE or degraded samples, please consult our technical team |
| Minimum input | 5 µg HMW gDNA (recommended 10–20 µg) | High-molecular-weight DNA (≥30 kb modal length) strongly recommended |
| DNA quality | OD260/280 1.8–2.0; OD260/230 ≥ 2.0; no visible degradation | Assessed by pulsed-field gel electrophoresis and TapeStation |
| Coverage | 10–30× HiFi (depending on package) | Higher coverage recommended for SV/TR discovery and haplotype phasing |
| Data transfer | Hard drive, FTP, or cloud transfer | Secure transfer arrangements made per project |
For projects involving re-sequencing of previously sequenced samples, we accept pre-extracted gDNA or sequencing-ready libraries. Sample shipment guidelines are available on our Sample Submission page.
Danzi MC, Xu IRL, Fazal S, Dolzhenko E, Pellerin D, Weisburd B, et al. (All of Us Research Program Long Read Working Group). bioRxiv 2025. DOI: 10.1101/2025.01.06.631535. (CC BY 4.0)
Tandem repeats (TRs) constitute some of the most polymorphic regions of the human genome and are implicated in over 60 known repeat expansion disorders. However, the full spectrum of human TR variation — including the allele sequences, length distributions, and population frequencies — has remained poorly characterised because short-read sequencing cannot resolve repetitive regions at nucleotide resolution. The All of Us Research Program generated PacBio HiFi whole-genome sequencing data from 1,027 participants, creating an opportunity to profile TR variation at an unprecedented scale.
The study used PacBio HiFi sequencing at ~8× coverage across 1,027 genomes from the All of Us cohort. TR genotyping was performed using the Tandem Repeat Genotyping Tool (TRGT), which leverages HiFi read accuracy to determine repeat length, motif sequence, and motif purity at each of 1.7+ million TR loci genome-wide. The resulting database comprises over 3.6 billion TR allele sequences, making it the largest publicly available HiFi-based TR resource to date.
Figure adapted from Danzi et al. 2025 (CC BY 4.0). Tandem repeat allele profiling across 1,027 HiFi genomes reveals population-level repeat purity, length distributions, and known pathogenic repeat loci as the most variable loci genome-wide.
This study demonstrates that population-scale HiFi TR genotyping is technically robust and biologically informative. The 3.6-billion-allele TR resource provides a critical reference for future repeat-disease association studies and establishes a framework for integrating TR variation into population genomics pipelines. It underscores the value of HiFi sequencing as a platform for comprehensive variant discovery in large cohorts.
Coverage requirements depend on your research objectives. For SV and TR discovery, 10–15× HiFi is sufficient for most population-scale studies, as demonstrated by the All of Us programme (8×). For comprehensive methylation profiling and haplotype phasing, we recommend 15–20×. For multi-omics integration (e.g., combining variant discovery with pangenome analysis and QTL mapping), 20–30× provides optimal power for downstream association analyses.
HiFi sequencing (>99.9% single-read accuracy) is better suited for applications requiring nucleotide-level precision across many samples — such as TR genotyping, SNV calling, and CNV copy-number estimation — because each read is consensus-corrected during sequencing. ONT sequencing offers longer reads (100+ kb) that can span the largest repetitive regions, making it complementary for closing assembly gaps and resolving complex SVs. For population-scale multi-omics analysis where consistent variant accuracy across hundreds or thousands of samples is critical, HiFi is the recommended primary platform.
Yes. Our Advanced analysis pipeline is designed for multi-omics integration. If you already have short-read WGS, RNA-seq, bisulfite sequencing, or proteomics data from the same cohort, we can integrate these with the HiFi variant call sets for eQTL, xQTL, methylation QTL, and multi-omics association analyses. This hybrid approach leverages the strengths of each data type — short-read depth and cost efficiency with HiFi's resolution of otherwise invisible variant classes.
The PacBio Revio platform produces approximately 200–250 Gb of HiFi data per SMRT Cell 25M run, equivalent to ~15 human genomes at 15× coverage per flow cell. For large cohorts (>500 samples), we batch samples across multiple Revio runs with standardised QC checkpoints between batches. Cohort sizes from 50 to several thousand samples are feasible, with project timelines scaled accordingly. Please contact our team to discuss your specific cohort size and project timeline.
Our standard pipeline uses GRCh38 (with ALT contigs) and the T2T-CHM13 reference genome. Where relevant to the study population, we can also align to additional reference panels or genome builds. Pangenome graph-based alignment (using minigraph or vg) is available in the Advanced module to reduce reference bias when working with diverse populations.
Our demo package includes representative deliverables from a pilot cohort HiFi WGS project:
1. Multi-variant call set browser view — SNV, SV, TR, CNV, and methylation tracks for a single sample displayed in IGV.
2. Cohort-level association summary — Manhattan plot, QQ plot, and TR allele distribution across the pilot cohort.
3. TR-xQTL results — example TR-phenotype association with locus zoom plot and repeat motif visualisation.
4. TWAS and EWAS results — transcriptome-wide and epigenome-wide association signals with gene-set enrichment and pathway analysis.
5. Multi-omics integration network — circos plot showing interactions between TRs, SNPs, gene expression, DNA methylation, and phenotype across QTL types.
Representative deliverable format: multi-variant association summary, TR allele distribution, and population-level QC metrics from a cohort HiFi WGS project.
Population-scale long-read projects require careful experimental design. Our scientific team — experienced in large-cohort HiFi sequencing, multi-variant calling, and multi-omics integration — can help you scope the optimal coverage, analysis modules, and project timeline for your specific research goals.
Submit an inquiry → or contact us at contact@cd-genomics.com | (631) 259-7705
References
Related Services