![]()
Human genomic regions that matter most for drug response and immune function — CYP2D6, CYP2C19, HLA-A/B/C, DRB1 — are also some of the most complex loci in the genome. When short-read NGS is applied to these regions, variant phasing collapses at repetitive elements, copy number variations (CNVs) remain invisible, and hybrid gene configurations are simply missed. For researchers who need definitive answers at these loci, the technology has to match the biology.
CD Genomics provides a dedicated human long-amplicon sequencing service built on PacBio HiFi circular consensus sequencing (CCS). Our approach delivers full-length, phased amplicon sequences spanning 500 bp to 20 kb with >99% single-molecule read accuracy — covering entire pharmacogenes, HLA loci, and immune receptor genes in a single read, without the assembly ambiguities that plague short-read amplicon studies.
We design the primers, run the PCR, handle the library preparation, operate the PacBio Sequel and Revio systems, and deliver CCS-generated consensus sequences with phased variant calls. The service supports pharmacogenomics, immunogenetics, immune repertoire profiling, rare disease variant confirmation, and population-scale targeted screening. If you are investigating a genomic region where short reads give you ambiguous or incomplete results, our long-amplicon workflow is designed to resolve it.
Short-read amplicon sequencing (150–300 bp paired-end) has an inherent limitation: it cannot physically connect two SNPs that are more than a few hundred bases apart. In genes like CYP2D6, where clinically relevant star alleles are defined by combinations of variants distributed across 4.4 kb of genomic sequence, short reads force computational phasing — statistical inference that is often wrong when rare or novel variants are present. The same problem exists in HLA genes, where phased exon combinations define 4-field typing resolution, and in antibody V(D)J recombination analysis, where full-length variable regions must be read in a single molecule to correctly identify clonotypes.
PacBio HiFi sequencing solves the phasing problem by design. Each SMRTbell library molecule is a closed circular template that the polymerase reads multiple times. The instrument generates a circular consensus sequence (CCS) from these multiple passes, producing a single high-accuracy read that represents the entire original amplicon molecule — including all its variants, in their true cis/trans arrangement. For a 5 kb CYP2D6 amplicon, this means a researcher receives a single read with all SNPs, INDELs, and CNV breakpoints correctly phased, exactly as they exist on each chromosome.
Our human long-amplicon sequencing service covers:
For researchers who need broader genomic coverage beyond targeted loci, our Human Whole Genome Sequencing service provides PacBio HiFi-based complete genome analysis with structural variant detection across the entire genome.
The core of our long-amplicon service is PacBio's HiFi sequencing, which generates circular consensus sequences through the following mechanism: each double-stranded DNA amplicon is ligated with hairpin SMRTbell adapters, creating a closed circular template. During sequencing on the PacBio Sequel II or Revio system, a DNA polymerase repeatedly traverses this circular template, generating multiple subreads of the same molecule. The instrument's onboard algorithms then align these subreads and compute a single high-accuracy consensus sequence — the HiFi read.
Because errors in individual subreads are random, the consensus calculation effectively eliminates them. A molecule read 10 times with 90% raw accuracy per pass yields a consensus accuracy exceeding 99.9%. For amplicons up to 12 kb, the polymerase can complete enough passes to achieve QV30+ (99.9% accuracy). For longer amplicons (12–20 kb), the number of passes per molecule decreases, but the consensus accuracy remains above 99%.
Workflow of human long-amplicon sequencing, integrating long PCR amplification, SMRTbell library preparation, PacBio CCS sequencing, and variant analysis with phased haplotypes.
The table below summarizes the key differences between the three approaches for targeted human genomic analysis:
| Feature | PacBio HiFi CCS (Our Service) | NGS Amplicon (Illumina) | Sanger Sequencing |
| Read Length | 500 bp–20 kb (full amplicon) | 150–300 bp (paired-end) | 800–1,000 bp per reaction |
| Single-Read Accuracy | >99% (QV30+) | >99.9% | >99.99% |
| Variant Phasing | Direct, per-molecule | Statistical imputation only | Not phased (single allele per reaction) |
| CNV Detection | Direct from read depth + breakpoint structure | Indirect (read depth alone) | Not applicable |
| Homopolymer Accuracy | High (CCS error correction) | Low (systematic errors) | High |
| Throughput per Run | Thousands of amplicons | Millions of amplicons | One amplicon per reaction |
| Novel Allele Discovery | Full-length de novo | Limited (assembly gaps) | Yes, but one amplicon at a time |
For researchers primarily evaluating platform options, our PacBio SMRT Sequencing Technology page provides a broader overview of the platform's capabilities across genomics, transcriptomics, and epigenetics applications.
Pharmacogenomics is one of the strongest use cases for long-amplicon sequencing because many pharmacogenes defy analysis by standard genotyping methods. Genes like CYP2D6 contain frequent CNVs (full-gene deletions, duplications, and multiplications), hybrid configurations with the neighboring pseudogene CYP2D7, and complex structural rearrangements — all of which directly affect enzyme activity and drug metabolism phenotype.
qPCR can estimate CYP2D6 copy number but cannot phase SNPs — it tells you how many copies exist, not which star alleles they carry. NGS short-read sequencing can identify individual SNPs but cannot determine whether two variants are on the same allele (cis) or different alleles (trans), and it systematically misses large structural variants because the reads are too short to span breakpoint junctions. Hybrid genes like CYP2D6*68+*4, where part of the gene is replaced by CYP2D7 sequence, are invisible to both approaches unless specifically targeted — yet they have profound functional consequences.
Our long-amplicon approach spans the entire CYP2D6 locus in 4–5 overlapping amplicons of 3–6 kb each, covering the promoter region, all 9 exons, all introns, and the 3' UTR. Each amplicon is sequenced as a single HiFi read, preserving the true arrangement of variants. The result is a definitive star-allele call — not a probabilistic imputation — that includes CNV status, hybrid gene detection, and novel variant identification.
The same approach extends to other pharmacogenes with complex architectures: CYP2C19 (structural variants affecting *2 and *17 alleles), NAT2 (haplotype-defining SNPs spread across 870 bp), TPMT (rare loss-of-function alleles), and UGT1A1 (promoter TA-repeat polymorphism).
Our team also supports broader pharmacogenomics programs through pharmacogenomics research solutions that combine long-amplicon sequencing with population genetic analysis and functional annotation.
The human leukocyte antigen (HLA) region on chromosome 6 contains the most polymorphic genes in the human genome. HLA-A, HLA-B, HLA-C (Class I) and HLA-DRB1, HLA-DQB1 (Class II) each have thousands of documented alleles — and the differences between them often involve combinations of variants distributed across multiple exons. Achieving 4-field resolution (which distinguishes alleles by all synonymous and non-synonymous coding differences) requires full-length, phased sequences across all exons and introns of each gene.
Short-read NGS (150–300 bp) cannot span entire HLA exons in a single read, let alone connect exon 2 to exon 3 for Class I genes or exon 2 to exon 3 for Class II. The result is ambiguous typing: when two variant combinations produce the same exon-level allele assignments in a given individual, short-read data cannot distinguish them. This ambiguity rate increases with sample diversity — populations with many novel or rare alleles, typical of underrepresented ancestry groups, produce more ambiguous calls.
Our HLA long-amplicon service generates PCR products of 3–5 kb that span entire HLA genes from the 5' UTR through the 3' UTR. For Class I genes (HLA-A, -B, -C), a single amplicon covers all exons and introns. For Class II genes (HLA-DRB1, -DQB1), amplicons are designed to cover the polymorphic exon 2 region plus flanking introns. Each amplicon is sequenced with PacBio HiFi CCS, producing a single phased read per allele.
The output includes 4-field HLA typing with phased haplotypes, novel allele identification (with supporting read evidence for submission to the IPD-IMGT/HLA database), and copy number assessment. For researchers focused specifically on HLA analysis, our dedicated Human HLA Typing service provides more detailed platform options including ONT-based approaches.
Full-length B-cell receptor (BCR) and T-cell receptor (TCR) repertoire sequencing requires complete coverage of the V(D)J variable region — typically 400–600 bp — plus sufficient flanking sequence to identify the gene segments used. Short-read NGS (2 × 300 bp) straddles the edge of this requirement, and for many receptor configurations, the CDR3 region alone consumes most of the read length, leaving framework regions uncovered.
PacBio HiFi CCS reads spanning 500–1,000 bp capture the entire V(D)J region in a single read, plus the constant region beginning, with error-corrected accuracy that matches or exceeds NGS quality. This combination — full-length coverage plus high accuracy — enables true clonotype identification, somatic hypermutation (SHM) analysis across the entire variable domain, and paired heavy/light chain reconstruction when combined with single-cell barcoding.
Full-length BCR repertoire analysis is particularly valuable for antibody discovery programs, where identifying rare clonotypes with specific CDR3 sequences and SHM patterns can lead to therapeutic candidates. For vaccine research, tracking longitudinal changes in BCR and TCR repertoires following immunization reveals clonal expansion dynamics, affinity maturation trajectories, and class-switch recombination events that short-read approaches may miss.
We recommend researchers also explore our Full-Length TCR/BCR Repertoire Profiling service for a comprehensive overview of our repertoire analysis capabilities, including single-cell and bulk approaches.
Beyond the primary pharmacogenomics and immunogenetics applications, long-amplicon sequencing addresses several other research scenarios where targeted long-read coverage makes a decisive difference.
When WES or WGS identifies a candidate variant in a gene that contains pseudogenes (e.g., PKD1, GBA, SMN1/SMN2), segmental duplications, or GC-rich repetitive elements, short-read validation often fails because reads map ambiguously or polymerase errors accumulate at homopolymer runs. Our long-amplicon service can target the specific locus with custom primers for a 2–15 kb amplicon, confirming the variant with full-length reads that span the problematic region and its unique flanking sequences.
Biobanks and population genetics consortia increasingly need to genotype pharmacogenes at scale — not just common tag SNPs, but complete star alleles including rare and novel variants. Our multiplex barcoding strategy allows hundreds of amplicons from dozens of individuals to be pooled in a single SMRT Cell, driving per-sample costs below the threshold where comprehensive pharmacogene typing becomes feasible at population scale. We provide custom primer panel design, barcoding optimization, and batch-level QC reporting.
Researchers working with structural variants more broadly may also benefit from our Human Genome Structural Variation Detection service, which covers whole-genome SV discovery using long-read sequencing.
Our long-amplicon sequencing service follows a structured workflow designed to minimize iteration and deliver analysis-ready data:
The bioinformatics pipeline for long-amplicon data follows three stages:
Stage 1 — CCS generation: Raw sequencing subreads from each SMRTbell molecule are aligned to generate a circular consensus sequence. Quality filtering removes CCS reads with fewer than the minimum number of passes or with consensus accuracy below QV20.
Stage 2 — Long amplicon analysis (LAA): CCS reads are aligned to the reference sequence of each target locus using minimap2. For each amplicon, the alignment identifies the boundaries between the target sequence and the SMRTbell adapter, trims adapters, and orients reads. Duplicate reads (same start/end coordinates) are collapsed to avoid PCR amplification bias in downstream variant calling.
Stage 3 — Variant calling and phasing: SNPs and short INDELs are called using DeepVariant or Clair3 with CCS-specific models. Structural variants and CNVs are detected from within-read evidence (for SVs within the amplicon span) or read-depth analysis (for whole-locus CNVs). Variants within each read are inherently phased — the final output is a VCF file with per-haplotype variant calls.
Standard deliverables include:
Our service accepts genomic DNA, biological samples for DNA extraction, and pre-amplified PCR products. The table below summarizes the minimum and recommended input amounts for each sample type.
| Sample Type | Recommended Amount | Minimum Amount | Quality Requirement |
| Genomic DNA | ≥500 ng | 200 ng | A260/A280 1.8–2.0, no visible degradation |
| Whole Blood | 2–5 mL | 1 mL | EDTA tube, shipped cold (4°C) |
| Saliva | 2 mL | 1 mL | Oragene or equivalent DNA collection kit |
| Cell Pellet | 1×10⁶ cells | 5×10⁵ cells | PBS washed, snap-frozen, shipped on dry ice |
| Tissue | 25 mg | 10 mg | Snap-frozen or RNAlater-preserved, shipped on dry ice |
| Purified PCR Product | ≥100 ng | 50 ng | Single band on gel, verified concentration |
Important notes:
CYP2D6 metabolizes approximately 25% of commonly prescribed drugs, and its genetic polymorphism has direct consequences for drug efficacy and toxicity. However, accurate CYP2D6 allele typing has been notoriously difficult because the locus features frequent CNVs (deletions, duplications, and multiplications of the entire gene), hybrid gene configurations with the neighboring CYP2D7 pseudogene, and complex structural rearrangements. Traditional SNP arrays, qPCR copy number assays, and short-read NGS each capture only part of this variation.
Charnaud and colleagues (2022) developed PLASTER — a comprehensive bioinformatics pipeline for allele typing from PacBio long-amplicon data — and applied it to a large-scale validation cohort. Long PCR amplicons spanning the full CYP2D6 locus were generated from genomic DNA of 377 Solomon Islanders and sequenced on the PacBio SMRT platform. The PLASTER pipeline performed CCS generation, read alignment, variant calling, CNV detection from read depth, and breakpoint analysis for hybrid gene detection.
The analysis identified 25 distinct CYP2D6 alleles across the cohort, including 8 novel alleles not previously catalogued in the PharmGKB database (Source: Charnaud et al., Communications Biology, 2022, Fig. 3). The pipeline accurately resolved copy number states — detecting homozygous and heterozygous deletions, single-copy normal, duplications, and multiplications — and identified CYP2D6/2D7 hybrid gene configurations through breakpoint clustering analysis. When classified by PharmGKB functional status, the allele frequency distribution revealed substantial population-specific variation, with several novel alleles predicted to affect protein function.
This study demonstrates that long-amplicon PacBio CCS sequencing, combined with a validated bioinformatics pipeline, delivers definitive CYP2D6 allele typing at population scale — resolving phasing, CNVs, and structural variants in a single workflow. The approach is directly generalizable to other complex pharmacogenes and targeted genomic regions where short-read methods produce ambiguous or incomplete results.
Our PacBio HiFi CCS platform reliably sequences amplicons from 500 bp to 20 kb in length, with optimal quality consistently achieved at 1–12 kb. For amplicons exceeding 12 kb, we optimize PCR conditions and library preparation to maintain high consensus accuracy throughout the full length.
PacBio HiFi CCS delivers >99% single-molecule read accuracy (QV30+), comparable to Sanger sequencing for individual bases. The key advantage over Sanger is throughput — we can simultaneously sequence thousands of amplicons across multiple samples with direct phasing, which Sanger cannot achieve. For projects requiring a small number of short amplicons, Sanger remains cost-effective; for multi-locus, multi-sample studies, HiFi CCS is the more efficient choice.
Yes. Because we sequence full-length amplicons, our bioinformatics pipeline can identify novel SNPs, INDELs, structural variants, and hybrid gene configurations without relying on reference databases. All novel variants are reported with supporting read evidence, including the number of CCS reads supporting each allele.
PacBio CCS generates circular consensus sequences by reading each template molecule multiple times. This repeated sampling effectively eliminates random homopolymer errors — the consensus accuracy exceeds 99% (QV30+). The systematic homopolymer errors that affect raw single-pass long reads are corrected in the consensus step, making HiFi reads suitable for applications requiring base-level precision.
Yes. We use barcoded SMRTbell adapters to multiplex hundreds of amplicons from multiple individuals per SMRT Cell. For population-scale studies, we design custom multiplexing strategies to balance read depth across targets and optimize per-sample cost. Our team will work with you to determine the optimal pooling strategy based on your target count, required coverage depth, and budget.
We recommend at least 200 ng of high-quality genomic DNA (A260/A280 ratio 1.8–2.0). For samples with limited material, our team can optimize down to 50 ng through adjusted PCR cycling conditions. We also accept diverse sample types — blood, saliva, tissue, and cell pellets — and perform DNA extraction as part of our service.
Because each PacBio HiFi read spans the entire amplicon in a single molecule, phasing is direct and unambiguous — there is no computational assembly required. Two heterozygous SNPs separated by 10 kb are read on the same molecule, yielding definitive haplotype assignment without statistical inference. This is a fundamental advantage over short-read sequencing, where phasing relies on population reference panels or parental genotype data.
Standard deliverables include demultiplexed CCS FASTQ files, aligned BAM files, variant calls in VCF format with phasing information, CNV analysis results, and a comprehensive analysis report. Custom bioinformatics — such as population genetic statistics, phylogenetic reconstruction, or integration with phenotype data — is available upon request.
We present representative PacBio HiFi CCS long-amplicon sequencing data from a CYP2D6 pharmacogene panel in a human genomic DNA sample. Three target regions across the CYP2D6 locus were amplified and sequenced.
The bar chart below shows the allele frequency distribution of CYP2D6 star alleles detected across a population cohort using our long-amplicon CCS pipeline, colored by PharmGKB functional status.
![]()
The histogram displays the length distribution of CCS reads generated from a 5 kb CYP2D6 amplicon, showing tight clustering around the expected amplicon size with minimal off-target or truncated reads, confirming high specificity of the PCR amplification and CCS generation steps.
An Integrative Genomics Viewer (IGV) screenshot demonstrates phased SNPs within a single HiFi CCS read spanning the full CYP2D6 amplicon. All heterozygous positions are correctly phased into two haplotypes, corresponding to the maternal and paternal alleles.
References
For Research Use Only. Not for use in diagnostic procedures.