De novo assembly, graph-based pangenome construction, structural variant discovery, and trait association analysis for population-scale genomics.
CD Genomics provides dedicated pangenome research solutions that combine long-read sequencing (PacBio HiFi and Oxford Nanopore) with comprehensive bioinformatics for population-scale genomic analysis. Unlike single-reference assembly projects, our pangenome service is designed for multi-sample cohort studies — from de novo genome assembly and graph-based pangenome construction through structural variant (SV), presence-absence variation (PAV), and copy-number variant (CNV) discovery, gene family classification, and trait association analysis. We support pangenome projects across plant, animal, microbial, and human species, with flexible project packages sized from pilot studies (5–10 samples) to large breeding and mechanism cohorts (200+ samples).
For Research Use Only. Not for use in diagnostic procedures, clinical decision-making, personal health assessment, or therapeutic decision-making.
A single reference genome cannot represent the full genetic diversity of a species. Pangenome analysis addresses this limitation by incorporating genomic information from multiple individuals, constructing a graph or matrix that captures both shared (core) and variable (dispensable/private) genomic content. However, the resolution and biological value of any pangenome depend entirely on the sequencing technology used to build it.
Short-read sequencing (<150–300 bp) has been the predominant approach for pangenome studies, but it has fundamental blind spots. Structural variants larger than the read length, tandem repeat expansions, segmental duplications, and insertion sequences (including transposable elements and endogenous viral elements) are poorly resolved or completely invisible to short-read methods. Presence-absence variation in repetitive regions — a major source of functional diversity in plants and animals — is systematically underestimated. Studies comparing short-read and long-read pangenomes consistently report that short-read approaches miss 50–70% of structural variants and substantially underestimate the dispensable genome fraction.
Long-read sequencing overcomes these limitations. PacBio HiFi reads (>99.9% accuracy, 15–25 kb average length) provide single-nucleotide resolution across the full length of SVs and repeats, enabling direct genotyping of insertions, deletions, inversions, and translocations. Oxford Nanopore ultralong reads (100+ kb) span the largest repetitive elements and complex structural rearrangements, providing contiguity for de novo assembly and gap closure. When combined, these platforms deliver pangenome assemblies and variant call sets that capture the true extent of genetic diversity within a population — including the large-effect variants most likely to underlie phenotypic variation, disease susceptibility, and adaptive traits.
Our pangenome service is structured to accommodate diverse study systems and research objectives. Below we outline four representative scenarios, each with characteristic sample numbers, data requirements, and analytical goals.
Major crops including rice, wheat, maize, soybean, cotton, and Brassica species have complex genomes with high repeat content (60–85%), extensive presence-absence variation, and large structural rearrangements that directly influence yield, disease resistance, and environmental adaptation. Long-read pangenome analysis in plants captures dispensable genes, resistance gene clusters, and TE-driven regulatory variation that short-read approaches systematically miss. Typical cohort: 10–200+ accessions covering global germplasm diversity.
Livestock species (pig, cattle, sheep, chicken, horse, goat) and aquaculture species (salmon, tilapia, shrimp, catfish) exhibit extensive structural variation affecting meat quality, growth rate, disease resistance, and reproductive traits. Pangenome analysis in animals enables identification of breed-specific SVs and PAVs, characterisation of gene presence-absence in immune and metabolic pathways, and association of structural variants with economically important phenotypes. Typical cohort: 10–100+ individuals across breeds.
Bacterial and fungal species are defined by their pangenomes: core genes shared across all strains and accessory genes that confer antibiotic resistance, virulence, host adaptation, and metabolic versatility. Long-read sequencing resolves the plasmid, phage, and genomic island content that drives accessory genome evolution — content that is largely invisible to short-read assembly due to repetitive mobile elements. Typical cohort: 20–500+ isolates or metagenome-assembled genomes.
The Human Pangenome Reference Consortium (HPRC), Chinese Pangenome Consortium (CPC), and Arab Pangenome Project have demonstrated that a single human reference genome (GRCh38 or T2T-CHM13) misses millions of bases of population-specific sequence and hundreds of thousands of structural variants per population. Human pangenome analysis enables unbiased SV discovery in understudied populations, improves short-read mapping rates in repetitive and divergent regions, and identifies disease-associated structural variants that are invisible to reference-based genotyping. Typical cohort: 10–1,000+ individuals.
Every pangenome project is built on a foundation of high-quality individual genome assemblies, from which we construct graph-based pangenomes, call structural variants, classify gene content, and perform trait association analyses. Our deliverable structure is designed to provide both raw data (for downstream custom analysis) and interpreted results (publication-ready figures and reports).
Our analysis pipeline is organised into two tiers — Basic and Advanced — allowing researchers to select the depth of analysis appropriate for their study objectives and budget. The Basic module covers core pangenome assembly and structural variant discovery; the Advanced module adds haplotype resolution, pangenome graph mining, and full trait association analysis.
| Analysis Feature | Basic | Advanced |
| De novo genome assembly (HiFi + ONT + Hi-C) | ✓ | ✓ |
| Assembly QC — completeness, contiguity, accuracy (BUSCO, QV, NG50) | ✓ | ✓ |
| Repeat annotation (EDTA, RepeatMasker) | ✓ | ✓ |
| Gene prediction & functional annotation (BRAKER/MAKER, InterPro, GO, KEGG) | ✓ | ✓ |
| Graph-based pangenome construction (Minigraph-Cactus / PGGB) | ✓ | ✓ |
| SV/PAV detection and genotyping across all samples | ✓ | ✓ |
| Gene family classification — core / dispensable / private | ✓ | ✓ |
| Phylogenetic analysis and population structure (SNP + SV-based) | ✓ | ✓ |
| Functional enrichment of core, dispensable, and private gene sets | ✓ | ✓ |
| T2T gap-free assembly of select genomes | — | ✓ |
| Haplotype-resolved (phased) assembly | — | ✓ |
| CNV detection and genotyping | — | ✓ |
| Pangenome graph visualisation and subgraph mining | — | ✓ |
| SV/PAV-based GWAS and trait association | — | ✓ |
| Selection sweep analysis (XP-CLR, iHS, Fst) and selective-sweep gene identification | — | ✓ |
| Candidate gene prioritisation and functional interpretation | — | ✓ |
| TE insertion polymorphism analysis | — | ✓ |
| Publication-ready figures (pangenome graph, SV landscape, GWAS Manhattan, selection sweep) | — | ✓ |
Project scoping depends on research objectives, genome complexity, and desired analytical depth. We recommend the following package configurations as starting points for discussion. Specific sample numbers, coverage levels, and analytical scope can be adjusted to match study requirements.
| Package | Recommended Samples | Sequencing per Sample | Analysis Module | Best Suited For |
| Pilot | 5–10 | 30–60× HiFi + 60–100× ONT + Hi-C | Basic | Proof-of-concept pangenome studies, initial SV landscape characterisation in a new species, method validation, grant proposal preliminary data |
| Population | 10–50 | 30–60× HiFi + 60–100× ONT + Hi-C | Basic or Advanced | Population-scale pangenome construction, core/dispensable genome definition, SV frequency cataloguing, phylogenetic and population structure analysis across diverse accessions |
| Breeding & Mechanism | 50–200+ | 15–30× HiFi (screening) + select high-depth for assembly | Advanced | SV/PAV-based GWAS and trait association, selection sweep detection, candidate gene discovery for breeding programmes, large-cohort association studies, multi-omics integration with transcriptomic or epigenomic data |
Reference benchmark. The above coverage recommendations and project designs are informed by published long-read pangenome studies including the 32-ecotype Arabidopsis thaliana pangenome built from PacBio HiFi data (Kang et al. 2023), the 53-individual Arab human pangenome using HiFi + ONT ultralong sequencing (Nassir et al. 2025), and the HPRC human pangenome draft (Liao et al. 2023). These studies demonstrate that long-read pangenome analysis at 10–60× HiFi coverage produces assemblies and variant call sets that substantially exceed the resolution achievable with short-read approaches.
High-molecular-weight (HMW) DNA extraction from tissue, blood, or cultured cells. Sample quality assessed by pulsed-field gel electrophoresis (PFE) and fluorometric quantification. Minimum modal fragment length: ≥30 kb for HiFi, ≥50 kb for ONT ultralong.
SMRTbell library preparation for PacBio HiFi (Revio or Sequel IIe) and/or ONT library preparation (ultralong protocol, R10.4.1 flow cells). Hi-C library preparation for chromosome-scale scaffolding where genome assembly is required.
Individual genome assembly for each sample using HiFi + ONT + Hi-C data. Assembly polishing, QC (BUSCO, QV, NG50), repeat annotation, gene prediction, and functional annotation. Haplotype-resolved assembly (Advanced module).
Overview of the pangenome analysis workflow from HMW DNA extraction and long-read sequencing through de novo assembly, graph-based pangenome construction, SV/PAV discovery, and trait association analysis.
Graph-based pangenome construction using Minigraph-Cactus or PGGB. SV/PAV/CNV detection and genotyping across all samples. Gene family classification (core / dispensable / private). Pangenome graph visualisation (Advanced: subgraph mining, locus-specific graph extraction).
SV/PAV-based GWAS, selection sweep analysis (XP-CLR, iHS, Fst), candidate gene prioritisation, and functional enrichment. Publication-ready figures: Manhattan plots, selection sweep landscapes, pangenome graph snapshots, SV frequency distributions, and gene category enrichment bar plots.
Comprehensive project report with assembly statistics, pangenome summary metrics, variant call sets (VCF), gene category lists, association results, publication-ready figures, and full data archiving on secure storage.
| Category | Requirement | Notes |
| Sample type | High-quality gDNA (blood, tissue, cultured cells, or high-quality gDNA stock) | For plants: etiolated seedlings or young leaf tissue recommended to minimise polyphenol and polysaccharide contamination |
| Minimum input | >5 µg (HiFi); >10 µg (ONT ultralong + HiFi combined); >1 µg (Hi-C) | HMW DNA extraction service available for challenging sample types |
| Purity | OD 260/280: 1.8–2.0; OD 260/230: >2.0 | No visible RNA contamination (RNase treatment included) |
| Fragment length | >30 kb modal (HiFi); >50 kb modal (ONT ultralong) | PFE or Femto Pulse QC prior to library construction |
| Sample format | Fresh tissue on dry ice, or gDNA in TE buffer on ice | Long-term storage at −80°C recommended |
| Category | Deliverable |
| Raw data | FASTQ files (HiFi, ONT, Hi-C), sequencing run summary reports (Q-score distribution, yield, read length) |
| Assemblies | Individual genome assemblies (FASTA), assembly QC reports (BUSCO, QV, NG50, L50, genome size, completeness), repeat annotation (GFF3), gene annotation (GFF3, protein/transcript FASTA) |
| Pangenome | Pangenome graph (GFA / VG format), core/dispensable/private gene lists (TSV), gene category functional enrichment results, pangenome summary statistics |
| Variant call sets | SV/PAV VCF across all samples, CNV calls (BED/seg), TE insertion polymorphism calls (VCF), variant frequency and distribution summary |
| Association & results | GWAS summary statistics (Advanced), selection sweep results, candidate gene lists with functional annotation (Advanced), publication-ready figures (Manhattan, QQ, selection sweep, pangenome graph, SV distribution) |
| Report | Project report (PDF) documenting methods, assembly and pangenome metrics, variant summary, association results, and interpretation |
To demonstrate the power of long-read pangenome analysis, we highlight the work of Kang et al. (2023), who constructed a graph-based pangenome of Arabidopsis thaliana from 32 ecotypes sequenced with PacBio HiFi long reads. This study, published in Nature Communications (CC BY 4.0), illustrates how long-read sequencing at population scale captures structural variation that is invisible to short-read methods and connects it to local adaptation.
Arabidopsis thaliana is the premier plant model system, but existing pangenome studies relied primarily on short-read data that could not resolve large structural variants or repeat-rich regions. The ecotypes selected for this study span the species' global range, including relict populations from the Qinghai-Tibet Plateau (Tibet-0), Italy, and Morocco — representing ancestral lineages — and postglacial expansion ecotypes from across Eurasia.
The authors generated de novo assemblies for all 32 ecotypes using PacBio HiFi long-read sequencing (mean coverage depth sufficient for high-quality assembly). A graph-based pangenome was constructed using the Minigraph-Cactus pipeline, producing a variation graph with 468,168 nodes and 649,692 edges spanning 243.27 Mb of genomic sequence. Structural variants were called and genotyped across all ecotypes, with transposable element associations analysed for each SV class.
Representative pangenome analysis results from Kang et al. (2023). Panels adapted from the original publication under CC BY 4.0. Left: Pangenome graph visualisation showing structural variant density across the A. thaliana genome. Right: PAV distribution across 32 ecotypes and identification of an HPCA1 promoter SV associated with alpine adaptation in the Tibet-0 ecotype.
This study demonstrates that long-read-based pangenome analysis uncovers structural variation at a resolution and scale that short-read methods cannot achieve. The identification of an adaptive SV directly linked to environmental adaptation illustrates the biological and practical value of population-scale long-read pangenome projects for understanding the genetic basis of phenotypic variation.
The number of samples depends on the research question. For a pilot pangenome characterisation of a new species, 5–10 samples can reveal the major axes of structural variation. For defining core vs. dispensable genome content and constructing a population-level pangenome, 20–50 samples across the species' diversity range is typical. For trait association and breeding applications, 100–200+ samples with phenotype data provides sufficient statistical power for SV-based GWAS. We can help you determine the optimal sample number based on your specific study design and desired statistical power.
For de novo genome assembly of the selected samples, we recommend 30–60× HiFi coverage and 60–100× ONT ultralong coverage, plus Hi-C for chromosome-scale scaffolding. For large-cohort pangenome projects (100+ samples), a two-tier strategy is cost-effective: sequence all samples at 15–30× HiFi for SV/PAV genotyping, and select 10–30 representative samples for deep-coverage (30–60×) assembly and high-quality pangenome construction. We can design a stratified sequencing strategy tailored to your budget and research goals.
Yes. If you have short-read WGS data from the same samples or population, we can integrate it with new long-read data in several ways: (1) short reads can be used for assembly polishing and base-level accuracy validation; (2) SNP genotypes from short-read data can be combined with SV/PAV genotypes from long-read data for joint association analysis; (3) transcriptomic short-read data (RNA-seq) can be integrated for expression QTL analysis linked to pangenome variation. Please discuss your existing data with our project scientists for the most efficient integration strategy.
Large genomes (e.g., wheat, ~16 Gb; pine, ~20–30 Gb) and polyploid genomes require tailored sequencing and assembly strategies. For large genomes, we optimise coverage depth and may recommend ONT ultralong as the primary platform for contiguity, supplemented with HiFi for accuracy. For polyploids, we use Hi-C or parental read-based phasing to resolve homoeologous chromosomes, and specialised polyploid-aware assemblers and SV callers. Our project scientists have experience with complex genomes across plant, animal, and fungal lineages. We can design a custom strategy for your species.
Below are representative output formats from a long-read pangenome project. These AI-generated illustrations demonstrate the types of results delivered with each project, including pangenome graph visualisation, SV landscape summary, and association results.
1. Pangenome graph visualisation showing structural variant distribution across assembled genomes and the core/dispensable/private gene classification summary.
2. SV/PAV/CNV landscape summary including variant type distribution, size spectrum, and population frequency spectrum across the study cohort.
3. Trait association results (Advanced module) — Manhattan plot from SV-based GWAS and selection sweep landscape across the genome.
Representative pangenome analysis deliverables: pangenome graph visualisation, SV/PAV distribution across samples, and GWAS/selection sweep results. AI-generated representative data — not from a specific study.
To discuss your pangenome project requirements, sample numbers, species-specific considerations, or to request a detailed quotation, please contact our scientific team.
References