Human Long-Amplicon Sequencing Service — Full-Length, Phased PacBio HiFi Amplicon Analysis

Human Long-Amplicon Sequencing Service — Full-Length, Phased PacBio HiFi Amplicon Analysis

Human long-amplicon sequencing PacBio HiFi CCS phased haplotypes illustration

Human genomic regions that matter most for drug response and immune function — CYP2D6, CYP2C19, HLA-A/B/C, DRB1 — are also some of the most complex loci in the genome. When short-read NGS is applied to these regions, variant phasing collapses at repetitive elements, copy number variations (CNVs) remain invisible, and hybrid gene configurations are simply missed. For researchers who need definitive answers at these loci, the technology has to match the biology.

CD Genomics provides a dedicated human long-amplicon sequencing service built on PacBio HiFi circular consensus sequencing (CCS). Our approach delivers full-length, phased amplicon sequences spanning 500 bp to 20 kb with >99% single-molecule read accuracy — covering entire pharmacogenes, HLA loci, and immune receptor genes in a single read, without the assembly ambiguities that plague short-read amplicon studies.

We design the primers, run the PCR, handle the library preparation, operate the PacBio Sequel and Revio systems, and deliver CCS-generated consensus sequences with phased variant calls. The service supports pharmacogenomics, immunogenetics, immune repertoire profiling, rare disease variant confirmation, and population-scale targeted screening. If you are investigating a genomic region where short reads give you ambiguous or incomplete results, our long-amplicon workflow is designed to resolve it.

Why Long Amplicon Sequencing for Human Genomics

The Phasing Problem in Pharmacogenes and Immune Loci

Short-read amplicon sequencing (150–300 bp paired-end) has an inherent limitation: it cannot physically connect two SNPs that are more than a few hundred bases apart. In genes like CYP2D6, where clinically relevant star alleles are defined by combinations of variants distributed across 4.4 kb of genomic sequence, short reads force computational phasing — statistical inference that is often wrong when rare or novel variants are present. The same problem exists in HLA genes, where phased exon combinations define 4-field typing resolution, and in antibody V(D)J recombination analysis, where full-length variable regions must be read in a single molecule to correctly identify clonotypes.

How PacBio HiFi CCS Changes Amplicon Sequencing

PacBio HiFi sequencing solves the phasing problem by design. Each SMRTbell library molecule is a closed circular template that the polymerase reads multiple times. The instrument generates a circular consensus sequence (CCS) from these multiple passes, producing a single high-accuracy read that represents the entire original amplicon molecule — including all its variants, in their true cis/trans arrangement. For a 5 kb CYP2D6 amplicon, this means a researcher receives a single read with all SNPs, INDELs, and CNV breakpoints correctly phased, exactly as they exist on each chromosome.

Service Scope and Core Capabilities

Our human long-amplicon sequencing service covers:

For researchers who need broader genomic coverage beyond targeted loci, our Human Whole Genome Sequencing service provides PacBio HiFi-based complete genome analysis with structural variant detection across the entire genome.

PacBio HiFi CCS Technology for Long Amplicon Analysis

Circular Consensus Sequencing: How Accuracy Is Achieved

The core of our long-amplicon service is PacBio's HiFi sequencing, which generates circular consensus sequences through the following mechanism: each double-stranded DNA amplicon is ligated with hairpin SMRTbell adapters, creating a closed circular template. During sequencing on the PacBio Sequel II or Revio system, a DNA polymerase repeatedly traverses this circular template, generating multiple subreads of the same molecule. The instrument's onboard algorithms then align these subreads and compute a single high-accuracy consensus sequence — the HiFi read.

Because errors in individual subreads are random, the consensus calculation effectively eliminates them. A molecule read 10 times with 90% raw accuracy per pass yields a consensus accuracy exceeding 99.9%. For amplicons up to 12 kb, the polymerase can complete enough passes to achieve QV30+ (99.9% accuracy). For longer amplicons (12–20 kb), the number of passes per molecule decreases, but the consensus accuracy remains above 99%.

Human long-amplicon sequencing workflow from DNA to CCS to phased variant analysisWorkflow of human long-amplicon sequencing, integrating long PCR amplification, SMRTbell library preparation, PacBio CCS sequencing, and variant analysis with phased haplotypes.

How HiFi CCS Compares to NGS Amplicon and Sanger

The table below summarizes the key differences between the three approaches for targeted human genomic analysis:

Feature PacBio HiFi CCS (Our Service) NGS Amplicon (Illumina) Sanger Sequencing
Read Length500 bp–20 kb (full amplicon)150–300 bp (paired-end)800–1,000 bp per reaction
Single-Read Accuracy>99% (QV30+)>99.9%>99.99%
Variant PhasingDirect, per-moleculeStatistical imputation onlyNot phased (single allele per reaction)
CNV DetectionDirect from read depth + breakpoint structureIndirect (read depth alone)Not applicable
Homopolymer AccuracyHigh (CCS error correction)Low (systematic errors)High
Throughput per RunThousands of ampliconsMillions of ampliconsOne amplicon per reaction
Novel Allele DiscoveryFull-length de novoLimited (assembly gaps)Yes, but one amplicon at a time

For researchers primarily evaluating platform options, our PacBio SMRT Sequencing Technology page provides a broader overview of the platform's capabilities across genomics, transcriptomics, and epigenetics applications.

Pharmacogene Typing — Resolve CYP2D6, CYP2C19 and Beyond

Pharmacogenomics is one of the strongest use cases for long-amplicon sequencing because many pharmacogenes defy analysis by standard genotyping methods. Genes like CYP2D6 contain frequent CNVs (full-gene deletions, duplications, and multiplications), hybrid configurations with the neighboring pseudogene CYP2D7, and complex structural rearrangements — all of which directly affect enzyme activity and drug metabolism phenotype.

Why NGS and qPCR Fail at CYP2D6 Typing

qPCR can estimate CYP2D6 copy number but cannot phase SNPs — it tells you how many copies exist, not which star alleles they carry. NGS short-read sequencing can identify individual SNPs but cannot determine whether two variants are on the same allele (cis) or different alleles (trans), and it systematically misses large structural variants because the reads are too short to span breakpoint junctions. Hybrid genes like CYP2D6*68+*4, where part of the gene is replaced by CYP2D7 sequence, are invisible to both approaches unless specifically targeted — yet they have profound functional consequences.

Full-Gene Coverage With Phased Haplotypes

Our long-amplicon approach spans the entire CYP2D6 locus in 4–5 overlapping amplicons of 3–6 kb each, covering the promoter region, all 9 exons, all introns, and the 3' UTR. Each amplicon is sequenced as a single HiFi read, preserving the true arrangement of variants. The result is a definitive star-allele call — not a probabilistic imputation — that includes CNV status, hybrid gene detection, and novel variant identification.

The same approach extends to other pharmacogenes with complex architectures: CYP2C19 (structural variants affecting *2 and *17 alleles), NAT2 (haplotype-defining SNPs spread across 870 bp), TPMT (rare loss-of-function alleles), and UGT1A1 (promoter TA-repeat polymorphism).

Our team also supports broader pharmacogenomics programs through pharmacogenomics research solutions that combine long-amplicon sequencing with population genetic analysis and functional annotation.

HLA Typing — 4-Field Resolution With Phased Haplotypes

The human leukocyte antigen (HLA) region on chromosome 6 contains the most polymorphic genes in the human genome. HLA-A, HLA-B, HLA-C (Class I) and HLA-DRB1, HLA-DQB1 (Class II) each have thousands of documented alleles — and the differences between them often involve combinations of variants distributed across multiple exons. Achieving 4-field resolution (which distinguishes alleles by all synonymous and non-synonymous coding differences) requires full-length, phased sequences across all exons and introns of each gene.

The Challenge of HLA Polymorphism for Short-Read Sequencing

Short-read NGS (150–300 bp) cannot span entire HLA exons in a single read, let alone connect exon 2 to exon 3 for Class I genes or exon 2 to exon 3 for Class II. The result is ambiguous typing: when two variant combinations produce the same exon-level allele assignments in a given individual, short-read data cannot distinguish them. This ambiguity rate increases with sample diversity — populations with many novel or rare alleles, typical of underrepresented ancestry groups, produce more ambiguous calls.

Full-Length HLA Amplicons

Our HLA long-amplicon service generates PCR products of 3–5 kb that span entire HLA genes from the 5' UTR through the 3' UTR. For Class I genes (HLA-A, -B, -C), a single amplicon covers all exons and introns. For Class II genes (HLA-DRB1, -DQB1), amplicons are designed to cover the polymorphic exon 2 region plus flanking introns. Each amplicon is sequenced with PacBio HiFi CCS, producing a single phased read per allele.

The output includes 4-field HLA typing with phased haplotypes, novel allele identification (with supporting read evidence for submission to the IPD-IMGT/HLA database), and copy number assessment. For researchers focused specifically on HLA analysis, our dedicated Human HLA Typing service provides more detailed platform options including ONT-based approaches.

Immune Repertoire Profiling — Full-Length BCR and TCR Analysis

Full-length B-cell receptor (BCR) and T-cell receptor (TCR) repertoire sequencing requires complete coverage of the V(D)J variable region — typically 400–600 bp — plus sufficient flanking sequence to identify the gene segments used. Short-read NGS (2 × 300 bp) straddles the edge of this requirement, and for many receptor configurations, the CDR3 region alone consumes most of the read length, leaving framework regions uncovered.

Full-Length Coverage With Error-Corrected Accuracy

PacBio HiFi CCS reads spanning 500–1,000 bp capture the entire V(D)J region in a single read, plus the constant region beginning, with error-corrected accuracy that matches or exceeds NGS quality. This combination — full-length coverage plus high accuracy — enables true clonotype identification, somatic hypermutation (SHM) analysis across the entire variable domain, and paired heavy/light chain reconstruction when combined with single-cell barcoding.

Applications in Antibody Discovery and Vaccine Research

Full-length BCR repertoire analysis is particularly valuable for antibody discovery programs, where identifying rare clonotypes with specific CDR3 sequences and SHM patterns can lead to therapeutic candidates. For vaccine research, tracking longitudinal changes in BCR and TCR repertoires following immunization reveals clonal expansion dynamics, affinity maturation trajectories, and class-switch recombination events that short-read approaches may miss.

We recommend researchers also explore our Full-Length TCR/BCR Repertoire Profiling service for a comprehensive overview of our repertoire analysis capabilities, including single-cell and bulk approaches.

Additional Applications — Rare Disease, Structural Variants, and Population Screening

Beyond the primary pharmacogenomics and immunogenetics applications, long-amplicon sequencing addresses several other research scenarios where targeted long-read coverage makes a decisive difference.

Rare Disease Variant Confirmation at Complex Loci

When WES or WGS identifies a candidate variant in a gene that contains pseudogenes (e.g., PKD1, GBA, SMN1/SMN2), segmental duplications, or GC-rich repetitive elements, short-read validation often fails because reads map ambiguously or polymerase errors accumulate at homopolymer runs. Our long-amplicon service can target the specific locus with custom primers for a 2–15 kb amplicon, confirming the variant with full-length reads that span the problematic region and its unique flanking sequences.

Population-Scale Pharmacogene Screening

Biobanks and population genetics consortia increasingly need to genotype pharmacogenes at scale — not just common tag SNPs, but complete star alleles including rare and novel variants. Our multiplex barcoding strategy allows hundreds of amplicons from dozens of individuals to be pooled in a single SMRT Cell, driving per-sample costs below the threshold where comprehensive pharmacogene typing becomes feasible at population scale. We provide custom primer panel design, barcoding optimization, and batch-level QC reporting.

Researchers working with structural variants more broadly may also benefit from our Human Genome Structural Variation Detection service, which covers whole-genome SV discovery using long-read sequencing.

Service Workflow and Bioinformatics Pipeline

Project Workflow: From Sample to Report

Our long-amplicon sequencing service follows a structured workflow designed to minimize iteration and deliver analysis-ready data:

  1. Consultation and primer design — We work with you to define target regions, design PCR primers for your loci of interest, and validate them in silico against the human reference genome to ensure specificity. For standard pharmacogenes and HLA loci, we maintain pre-validated primer panels.
  2. PCR amplification and QC — Long-range PCR is performed on your DNA samples using high-fidelity polymerase optimized for GC-rich and repetitive templates. Amplicons are verified by gel electrophoresis or Bioanalyzer for correct size and yield.
  3. SMRTbell library preparation — Amplicons are end-repaired, A-tailed, and ligated with SMRTbell hairpin adapters. Barcoded adapters are used for multiplexing. Library concentration and fragment size are confirmed by Qubit and Bioanalyzer.
  4. PacBio CCS sequencing — Libraries are loaded onto SMRT Cells and sequenced on the PacBio Sequel II or Revio system. Sequencing duration is adjusted to achieve target CCS coverage (typically 100–500× per amplicon).
  5. Bioinformatics analysis — Raw subreads are processed into CCS reads, demultiplexed, and aligned to the human reference genome or user-specified target sequences.
  6. Data delivery — Final results are delivered through our secure data portal.

Bioinformatics Pipeline: CCS → LAA → Variant Analysis

The bioinformatics pipeline for long-amplicon data follows three stages:

Stage 1 — CCS generation: Raw sequencing subreads from each SMRTbell molecule are aligned to generate a circular consensus sequence. Quality filtering removes CCS reads with fewer than the minimum number of passes or with consensus accuracy below QV20.

Stage 2 — Long amplicon analysis (LAA): CCS reads are aligned to the reference sequence of each target locus using minimap2. For each amplicon, the alignment identifies the boundaries between the target sequence and the SMRTbell adapter, trims adapters, and orients reads. Duplicate reads (same start/end coordinates) are collapsed to avoid PCR amplification bias in downstream variant calling.

Stage 3 — Variant calling and phasing: SNPs and short INDELs are called using DeepVariant or Clair3 with CCS-specific models. Structural variants and CNVs are detected from within-read evidence (for SVs within the amplicon span) or read-depth analysis (for whole-locus CNVs). Variants within each read are inherently phased — the final output is a VCF file with per-haplotype variant calls.

Standard deliverables include:

Sample Requirements

Our service accepts genomic DNA, biological samples for DNA extraction, and pre-amplified PCR products. The table below summarizes the minimum and recommended input amounts for each sample type.

Sample Type Recommended Amount Minimum Amount Quality Requirement
Genomic DNA≥500 ng200 ngA260/A280 1.8–2.0, no visible degradation
Whole Blood2–5 mL1 mLEDTA tube, shipped cold (4°C)
Saliva2 mL1 mLOragene or equivalent DNA collection kit
Cell Pellet1×10⁶ cells5×10⁵ cellsPBS washed, snap-frozen, shipped on dry ice
Tissue25 mg10 mgSnap-frozen or RNAlater-preserved, shipped on dry ice
Purified PCR Product≥100 ng50 ngSingle band on gel, verified concentration

Important notes:

  • CD Genomics provides end-to-end amplicon generation — clients may submit genomic DNA or biological samples, and our team handles primer design, PCR amplification, and QC.
  • Amplicon size range: 500 bp to 20 kb. Optimal quality is consistently achieved for amplicons between 1 kb and 12 kb.
  • For population-scale studies, our team designs custom multiplex barcoding strategies to pool hundreds of amplicons from multiple individuals in a single sequencing run, optimizing per-sample cost while maintaining sufficient read depth per locus.

Case Study: CYP2D6 Allele Typing in a Population Cohort

Background: The CYP2D6 Typing Challenge

CYP2D6 metabolizes approximately 25% of commonly prescribed drugs, and its genetic polymorphism has direct consequences for drug efficacy and toxicity. However, accurate CYP2D6 allele typing has been notoriously difficult because the locus features frequent CNVs (deletions, duplications, and multiplications of the entire gene), hybrid gene configurations with the neighboring CYP2D7 pseudogene, and complex structural rearrangements. Traditional SNP arrays, qPCR copy number assays, and short-read NGS each capture only part of this variation.

Methods: PLASTER Pipeline and PacBio CCS

Charnaud and colleagues (2022) developed PLASTER — a comprehensive bioinformatics pipeline for allele typing from PacBio long-amplicon data — and applied it to a large-scale validation cohort. Long PCR amplicons spanning the full CYP2D6 locus were generated from genomic DNA of 377 Solomon Islanders and sequenced on the PacBio SMRT platform. The PLASTER pipeline performed CCS generation, read alignment, variant calling, CNV detection from read depth, and breakpoint analysis for hybrid gene detection.

Results: 25 Alleles, 8 Novel, Full Copy Number Resolution

The analysis identified 25 distinct CYP2D6 alleles across the cohort, including 8 novel alleles not previously catalogued in the PharmGKB database (Source: Charnaud et al., Communications Biology, 2022, Fig. 3). The pipeline accurately resolved copy number states — detecting homozygous and heterozygous deletions, single-copy normal, duplications, and multiplications — and identified CYP2D6/2D7 hybrid gene configurations through breakpoint clustering analysis. When classified by PharmGKB functional status, the allele frequency distribution revealed substantial population-specific variation, with several novel alleles predicted to affect protein function.

Conclusions: Scalable Pharmacogene Typing Validated

This study demonstrates that long-amplicon PacBio CCS sequencing, combined with a validated bioinformatics pipeline, delivers definitive CYP2D6 allele typing at population scale — resolving phasing, CNVs, and structural variants in a single workflow. The approach is directly generalizable to other complex pharmacogenes and targeted genomic regions where short-read methods produce ambiguous or incomplete results.

FAQs

Our PacBio HiFi CCS platform reliably sequences amplicons from 500 bp to 20 kb in length, with optimal quality consistently achieved at 1–12 kb. For amplicons exceeding 12 kb, we optimize PCR conditions and library preparation to maintain high consensus accuracy throughout the full length.

PacBio HiFi CCS delivers >99% single-molecule read accuracy (QV30+), comparable to Sanger sequencing for individual bases. The key advantage over Sanger is throughput — we can simultaneously sequence thousands of amplicons across multiple samples with direct phasing, which Sanger cannot achieve. For projects requiring a small number of short amplicons, Sanger remains cost-effective; for multi-locus, multi-sample studies, HiFi CCS is the more efficient choice.

Yes. Because we sequence full-length amplicons, our bioinformatics pipeline can identify novel SNPs, INDELs, structural variants, and hybrid gene configurations without relying on reference databases. All novel variants are reported with supporting read evidence, including the number of CCS reads supporting each allele.

PacBio CCS generates circular consensus sequences by reading each template molecule multiple times. This repeated sampling effectively eliminates random homopolymer errors — the consensus accuracy exceeds 99% (QV30+). The systematic homopolymer errors that affect raw single-pass long reads are corrected in the consensus step, making HiFi reads suitable for applications requiring base-level precision.

Yes. We use barcoded SMRTbell adapters to multiplex hundreds of amplicons from multiple individuals per SMRT Cell. For population-scale studies, we design custom multiplexing strategies to balance read depth across targets and optimize per-sample cost. Our team will work with you to determine the optimal pooling strategy based on your target count, required coverage depth, and budget.

We recommend at least 200 ng of high-quality genomic DNA (A260/A280 ratio 1.8–2.0). For samples with limited material, our team can optimize down to 50 ng through adjusted PCR cycling conditions. We also accept diverse sample types — blood, saliva, tissue, and cell pellets — and perform DNA extraction as part of our service.

Because each PacBio HiFi read spans the entire amplicon in a single molecule, phasing is direct and unambiguous — there is no computational assembly required. Two heterozygous SNPs separated by 10 kb are read on the same molecule, yielding definitive haplotype assignment without statistical inference. This is a fundamental advantage over short-read sequencing, where phasing relies on population reference panels or parental genotype data.

Standard deliverables include demultiplexed CCS FASTQ files, aligned BAM files, variant calls in VCF format with phasing information, CNV analysis results, and a comprehensive analysis report. Custom bioinformatics — such as population genetic statistics, phylogenetic reconstruction, or integration with phenotype data — is available upon request.

Demo Results

We present representative PacBio HiFi CCS long-amplicon sequencing data from a CYP2D6 pharmacogene panel in a human genomic DNA sample. Three target regions across the CYP2D6 locus were amplified and sequenced.

1. CYP2D6 Allele Frequency Distribution

The bar chart below shows the allele frequency distribution of CYP2D6 star alleles detected across a population cohort using our long-amplicon CCS pipeline, colored by PharmGKB functional status.

CYP2D6 allele frequency distribution bar chart from long-amplicon CCS sequencing

2. CCS Read Length Distribution

The histogram displays the length distribution of CCS reads generated from a 5 kb CYP2D6 amplicon, showing tight clustering around the expected amplicon size with minimal off-target or truncated reads, confirming high specificity of the PCR amplification and CCS generation steps.

3. Phased Variant Call Across a Full-Length CYP2D6 Amplicon

An Integrative Genomics Viewer (IGV) screenshot demonstrates phased SNPs within a single HiFi CCS read spanning the full CYP2D6 amplicon. All heterozygous positions are correctly phased into two haplotypes, corresponding to the maternal and paternal alleles.

References

  1. Charnaud S, Munro JE, Semenec L, et al. PacBio long-read amplicon sequencing enables scalable high-resolution population allele typing of the complex CYP2D6 locus. Communications Biology, 2022, 5(1): 168.
  2. Buermans HPJ, Vossen RHAM, Anvar SY, et al. Flexible and scalable full-length CYP2D6 long amplicon PacBio sequencing. Human Mutation, 2017, 38(3): 310-316.
  3. Wenger AM, Peluso P, Rowell WJ, et al. Accurate circular consensus long-read sequencing improves variant detection and assembly of a human genome. Nature Biotechnology, 2019, 37(10): 1155-1162.

For Research Use Only. Not for use in diagnostic procedures.

Get Your Instant Quote