Population-Scale Pangenome Research Solutions Using Long-Read Sequencing

Population-Scale Pangenome Research Solutions Using Long-Read Sequencing

De novo assembly, graph-based pangenome construction, structural variant discovery, and trait association analysis for population-scale genomics.

CD Genomics provides dedicated pangenome research solutions that combine long-read sequencing (PacBio HiFi and Oxford Nanopore) with comprehensive bioinformatics for population-scale genomic analysis. Unlike single-reference assembly projects, our pangenome service is designed for multi-sample cohort studies — from de novo genome assembly and graph-based pangenome construction through structural variant (SV), presence-absence variation (PAV), and copy-number variant (CNV) discovery, gene family classification, and trait association analysis. We support pangenome projects across plant, animal, microbial, and human species, with flexible project packages sized from pilot studies (5–10 samples) to large breeding and mechanism cohorts (200+ samples).

What Makes This Service Different

For Research Use Only. Not for use in diagnostic procedures, clinical decision-making, personal health assessment, or therapeutic decision-making.

Why Long-Read Sequencing for Pangenome Analysis?

A single reference genome cannot represent the full genetic diversity of a species. Pangenome analysis addresses this limitation by incorporating genomic information from multiple individuals, constructing a graph or matrix that captures both shared (core) and variable (dispensable/private) genomic content. However, the resolution and biological value of any pangenome depend entirely on the sequencing technology used to build it.

Short-read sequencing (<150–300 bp) has been the predominant approach for pangenome studies, but it has fundamental blind spots. Structural variants larger than the read length, tandem repeat expansions, segmental duplications, and insertion sequences (including transposable elements and endogenous viral elements) are poorly resolved or completely invisible to short-read methods. Presence-absence variation in repetitive regions — a major source of functional diversity in plants and animals — is systematically underestimated. Studies comparing short-read and long-read pangenomes consistently report that short-read approaches miss 50–70% of structural variants and substantially underestimate the dispensable genome fraction.

Long-read sequencing overcomes these limitations. PacBio HiFi reads (>99.9% accuracy, 15–25 kb average length) provide single-nucleotide resolution across the full length of SVs and repeats, enabling direct genotyping of insertions, deletions, inversions, and translocations. Oxford Nanopore ultralong reads (100+ kb) span the largest repetitive elements and complex structural rearrangements, providing contiguity for de novo assembly and gap closure. When combined, these platforms deliver pangenome assemblies and variant call sets that capture the true extent of genetic diversity within a population — including the large-effect variants most likely to underlie phenotypic variation, disease susceptibility, and adaptive traits.

Applications Across Species — Four Research Scenarios

Our pangenome service is structured to accommodate diverse study systems and research objectives. Below we outline four representative scenarios, each with characteristic sample numbers, data requirements, and analytical goals.

Plant Pangenomes — Crop Breeding & Trait Discovery

Major crops including rice, wheat, maize, soybean, cotton, and Brassica species have complex genomes with high repeat content (60–85%), extensive presence-absence variation, and large structural rearrangements that directly influence yield, disease resistance, and environmental adaptation. Long-read pangenome analysis in plants captures dispensable genes, resistance gene clusters, and TE-driven regulatory variation that short-read approaches systematically miss. Typical cohort: 10–200+ accessions covering global germplasm diversity.

  • SV/PAV discovery for agronomic trait association
  • Core vs. dispensable gene identification across cultivars
  • Resistance gene (NLR, RLK) pangenome dynamics
  • TE insertion polymorphism and gene expression impact

Animal Pangenomes — Livestock Breeding & Evolution

Livestock species (pig, cattle, sheep, chicken, horse, goat) and aquaculture species (salmon, tilapia, shrimp, catfish) exhibit extensive structural variation affecting meat quality, growth rate, disease resistance, and reproductive traits. Pangenome analysis in animals enables identification of breed-specific SVs and PAVs, characterisation of gene presence-absence in immune and metabolic pathways, and association of structural variants with economically important phenotypes. Typical cohort: 10–100+ individuals across breeds.

  • Breed-specific SV/CNV discovery and genotyping
  • Pangenome-guided GWAS for production traits
  • Selection signature detection in domestic vs. wild populations
  • Haplotype-resolved structural variant analysis

Microbial & Pathogen Pangenomes — AMR, Virulence & Evolution

Bacterial and fungal species are defined by their pangenomes: core genes shared across all strains and accessory genes that confer antibiotic resistance, virulence, host adaptation, and metabolic versatility. Long-read sequencing resolves the plasmid, phage, and genomic island content that drives accessory genome evolution — content that is largely invisible to short-read assembly due to repetitive mobile elements. Typical cohort: 20–500+ isolates or metagenome-assembled genomes.

  • Accessory genome characterisation (plasmids, phages, genomic islands)
  • Antibiotic resistance gene (ARG) pangenome profiling
  • Virulence factor presence-absence across lineages
  • Phylogenetic pangenome analysis for outbreak tracking

Human Pangenomes — Structural Variation Discovery & Disease Association

The Human Pangenome Reference Consortium (HPRC), Chinese Pangenome Consortium (CPC), and Arab Pangenome Project have demonstrated that a single human reference genome (GRCh38 or T2T-CHM13) misses millions of bases of population-specific sequence and hundreds of thousands of structural variants per population. Human pangenome analysis enables unbiased SV discovery in understudied populations, improves short-read mapping rates in repetitive and divergent regions, and identifies disease-associated structural variants that are invisible to reference-based genotyping. Typical cohort: 10–1,000+ individuals.

  • Population-specific SV/PAV discovery and frequency estimation
  • Pangenome graph alignment for reduced reference bias
  • Structural variant association with complex disease
  • Gene duplication and deletion burden testing in cohorts

What We Deliver — From Assemblies to Biological Insight

Every pangenome project is built on a foundation of high-quality individual genome assemblies, from which we construct graph-based pangenomes, call structural variants, classify gene content, and perform trait association analyses. Our deliverable structure is designed to provide both raw data (for downstream custom analysis) and interpreted results (publication-ready figures and reports).

Individual Genome Assemblies

  • De novo assembly using PacBio HiFi + ONT ultralong + Hi-C data, with contiguity targets of N50 >10 Mb (plant/animal) or T2T-complete (microbial/small genomes)
  • Assembly QC — completeness (BUSCO), contiguity (NG50, L50), base accuracy (QV >40), and coverage assessment
  • Genome annotation — repeat annotation (EDTA, RepeatMasker), gene prediction (BRAKER, MAKER), and functional annotation (InterPro, GO, KEGG)
  • Haplotype-resolved assembly (Advanced module) — phased diploid assemblies with Hi-C or parental read-based phasing

Pangenome Construction & Analysis

  • Graph-based pangenome — built with Minigraph-Cactus, PanGenome Graph Builder (PGGB), or minigraph, producing a variation graph with embedded assembly paths
  • SV/PAV/CNV discovery — structural variant calling and genotyping across all samples (deletions, insertions, inversions, duplications, translocations, mobile element insertions)
  • Gene family classification — orthologous gene clustering across assemblies, with core (all samples), dispensable (<100%), and private (single sample) gene categorisation and functional enrichment per category
  • Trait association — SV/PAV-based GWAS, candidate gene identification, selection sweep analysis (XP-CLR, iHS, Fst), and functional enrichment (GO, KEGG, Reactome)

Available Analysis Modules

Our analysis pipeline is organised into two tiers — Basic and Advanced — allowing researchers to select the depth of analysis appropriate for their study objectives and budget. The Basic module covers core pangenome assembly and structural variant discovery; the Advanced module adds haplotype resolution, pangenome graph mining, and full trait association analysis.

Analysis Feature Basic Advanced
De novo genome assembly (HiFi + ONT + Hi-C)
Assembly QC — completeness, contiguity, accuracy (BUSCO, QV, NG50)
Repeat annotation (EDTA, RepeatMasker)
Gene prediction & functional annotation (BRAKER/MAKER, InterPro, GO, KEGG)
Graph-based pangenome construction (Minigraph-Cactus / PGGB)
SV/PAV detection and genotyping across all samples
Gene family classification — core / dispensable / private
Phylogenetic analysis and population structure (SNP + SV-based)
Functional enrichment of core, dispensable, and private gene sets
T2T gap-free assembly of select genomes
Haplotype-resolved (phased) assembly
CNV detection and genotyping
Pangenome graph visualisation and subgraph mining
SV/PAV-based GWAS and trait association
Selection sweep analysis (XP-CLR, iHS, Fst) and selective-sweep gene identification
Candidate gene prioritisation and functional interpretation
TE insertion polymorphism analysis
Publication-ready figures (pangenome graph, SV landscape, GWAS Manhattan, selection sweep)

Recommended Project Packages

Project scoping depends on research objectives, genome complexity, and desired analytical depth. We recommend the following package configurations as starting points for discussion. Specific sample numbers, coverage levels, and analytical scope can be adjusted to match study requirements.

Package Recommended Samples Sequencing per Sample Analysis Module Best Suited For
Pilot 5–10 30–60× HiFi + 60–100× ONT + Hi-C Basic Proof-of-concept pangenome studies, initial SV landscape characterisation in a new species, method validation, grant proposal preliminary data
Population 10–50 30–60× HiFi + 60–100× ONT + Hi-C Basic or Advanced Population-scale pangenome construction, core/dispensable genome definition, SV frequency cataloguing, phylogenetic and population structure analysis across diverse accessions
Breeding & Mechanism 50–200+ 15–30× HiFi (screening) + select high-depth for assembly Advanced SV/PAV-based GWAS and trait association, selection sweep detection, candidate gene discovery for breeding programmes, large-cohort association studies, multi-omics integration with transcriptomic or epigenomic data

Reference benchmark. The above coverage recommendations and project designs are informed by published long-read pangenome studies including the 32-ecotype Arabidopsis thaliana pangenome built from PacBio HiFi data (Kang et al. 2023), the 53-individual Arab human pangenome using HiFi + ONT ultralong sequencing (Nassir et al. 2025), and the HPRC human pangenome draft (Liao et al. 2023). These studies demonstrate that long-read pangenome analysis at 10–60× HiFi coverage produces assemblies and variant call sets that substantially exceed the resolution achievable with short-read approaches.

Workflow — From Sample to Pangenome Report

1. Sample Preparation & QC

High-molecular-weight (HMW) DNA extraction from tissue, blood, or cultured cells. Sample quality assessed by pulsed-field gel electrophoresis (PFE) and fluorometric quantification. Minimum modal fragment length: ≥30 kb for HiFi, ≥50 kb for ONT ultralong.

2. Library Construction & Sequencing

SMRTbell library preparation for PacBio HiFi (Revio or Sequel IIe) and/or ONT library preparation (ultralong protocol, R10.4.1 flow cells). Hi-C library preparation for chromosome-scale scaffolding where genome assembly is required.

3. De Novo Assembly & Annotation

Individual genome assembly for each sample using HiFi + ONT + Hi-C data. Assembly polishing, QC (BUSCO, QV, NG50), repeat annotation, gene prediction, and functional annotation. Haplotype-resolved assembly (Advanced module).

Overview of the pangenome analysis workflow from HMW DNA extraction and long-read sequencing through de novo assembly, graph-based pangenome construction, SV/PAV discovery, and trait association analysis.

4. Pangenome Construction & Variant Discovery

Graph-based pangenome construction using Minigraph-Cactus or PGGB. SV/PAV/CNV detection and genotyping across all samples. Gene family classification (core / dispensable / private). Pangenome graph visualisation (Advanced: subgraph mining, locus-specific graph extraction).

5. Trait Association & Biological Interpretation

SV/PAV-based GWAS, selection sweep analysis (XP-CLR, iHS, Fst), candidate gene prioritisation, and functional enrichment. Publication-ready figures: Manhattan plots, selection sweep landscapes, pangenome graph snapshots, SV frequency distributions, and gene category enrichment bar plots.

6. Report & Data Delivery

Comprehensive project report with assembly statistics, pangenome summary metrics, variant call sets (VCF), gene category lists, association results, publication-ready figures, and full data archiving on secure storage.

Sample Requirements

Category Requirement Notes
Sample type High-quality gDNA (blood, tissue, cultured cells, or high-quality gDNA stock) For plants: etiolated seedlings or young leaf tissue recommended to minimise polyphenol and polysaccharide contamination
Minimum input >5 µg (HiFi); >10 µg (ONT ultralong + HiFi combined); >1 µg (Hi-C) HMW DNA extraction service available for challenging sample types
Purity OD 260/280: 1.8–2.0; OD 260/230: >2.0 No visible RNA contamination (RNase treatment included)
Fragment length >30 kb modal (HiFi); >50 kb modal (ONT ultralong) PFE or Femto Pulse QC prior to library construction
Sample format Fresh tissue on dry ice, or gDNA in TE buffer on ice Long-term storage at −80°C recommended

Deliverables

Category Deliverable
Raw data FASTQ files (HiFi, ONT, Hi-C), sequencing run summary reports (Q-score distribution, yield, read length)
Assemblies Individual genome assemblies (FASTA), assembly QC reports (BUSCO, QV, NG50, L50, genome size, completeness), repeat annotation (GFF3), gene annotation (GFF3, protein/transcript FASTA)
Pangenome Pangenome graph (GFA / VG format), core/dispensable/private gene lists (TSV), gene category functional enrichment results, pangenome summary statistics
Variant call sets SV/PAV VCF across all samples, CNV calls (BED/seg), TE insertion polymorphism calls (VCF), variant frequency and distribution summary
Association & results GWAS summary statistics (Advanced), selection sweep results, candidate gene lists with functional annotation (Advanced), publication-ready figures (Manhattan, QQ, selection sweep, pangenome graph, SV distribution)
Report Project report (PDF) documenting methods, assembly and pangenome metrics, variant summary, association results, and interpretation

Case Study: A Graph-Based Pangenome of 32 Arabidopsis thaliana Ecotypes Using PacBio HiFi Sequencing

To demonstrate the power of long-read pangenome analysis, we highlight the work of Kang et al. (2023), who constructed a graph-based pangenome of Arabidopsis thaliana from 32 ecotypes sequenced with PacBio HiFi long reads. This study, published in Nature Communications (CC BY 4.0), illustrates how long-read sequencing at population scale captures structural variation that is invisible to short-read methods and connects it to local adaptation.

1. Background

Arabidopsis thaliana is the premier plant model system, but existing pangenome studies relied primarily on short-read data that could not resolve large structural variants or repeat-rich regions. The ecotypes selected for this study span the species' global range, including relict populations from the Qinghai-Tibet Plateau (Tibet-0), Italy, and Morocco — representing ancestral lineages — and postglacial expansion ecotypes from across Eurasia.

2. Methods

The authors generated de novo assemblies for all 32 ecotypes using PacBio HiFi long-read sequencing (mean coverage depth sufficient for high-quality assembly). A graph-based pangenome was constructed using the Minigraph-Cactus pipeline, producing a variation graph with 468,168 nodes and 649,692 edges spanning 243.27 Mb of genomic sequence. Structural variants were called and genotyped across all ecotypes, with transposable element associations analysed for each SV class.

3. Results

Representative pangenome analysis results from Kang et al. (2023). Panels adapted from the original publication under CC BY 4.0. Left: Pangenome graph visualisation showing structural variant density across the A. thaliana genome. Right: PAV distribution across 32 ecotypes and identification of an HPCA1 promoter SV associated with alpine adaptation in the Tibet-0 ecotype.

Key Findings

4. Conclusions

This study demonstrates that long-read-based pangenome analysis uncovers structural variation at a resolution and scale that short-read methods cannot achieve. The identification of an adaptive SV directly linked to environmental adaptation illustrates the biological and practical value of population-scale long-read pangenome projects for understanding the genetic basis of phenotypic variation.

FAQs

Demo & Inquiries

Below are representative output formats from a long-read pangenome project. These AI-generated illustrations demonstrate the types of results delivered with each project, including pangenome graph visualisation, SV landscape summary, and association results.

1. Pangenome graph visualisation showing structural variant distribution across assembled genomes and the core/dispensable/private gene classification summary.

2. SV/PAV/CNV landscape summary including variant type distribution, size spectrum, and population frequency spectrum across the study cohort.

3. Trait association results (Advanced module) — Manhattan plot from SV-based GWAS and selection sweep landscape across the genome.

Representative pangenome analysis deliverables: pangenome graph visualisation, SV/PAV distribution across samples, and GWAS/selection sweep results. AI-generated representative data — not from a specific study.

To discuss your pangenome project requirements, sample numbers, species-specific considerations, or to request a detailed quotation, please contact our scientific team.

References

  1. Kang M, Wu H, Liu H, et al. The pan-genome and local adaptation of Arabidopsis thaliana. Nature Communications. 2023;14:6259. doi:10.1038/s41467-023-42029-4 (CC BY 4.0)
  2. Nassir N, Almarri MA, Kumail M, et al. A draft UAE-based Arab pangenome reference. Nature Communications. 2025;16:6747. doi:10.1038/s41467-025-61645-w (CC BY 4.0)
  3. Liao WW, Asri M, Ebler J, et al. A draft human pangenome reference. Nature. 2023;617:312–324. doi:10.1038/s41586-023-05896-x (reference only)
Get Your Instant Quote