Whole Exome and Targeted Region Sequencing Services: From Human Genetics to Custom Panel Design

The human exome — the complete set of protein-coding exons across approximately 20,000 genes — represents roughly 1 percent of the genome but harbors an estimated 85 percent of known disease-causing variants. This asymmetric distribution of functional variation makes exome sequencing one of the most efficient experimental designs in genomics: target the small fraction of the genome where the vast majority of clinically interpretable information resides, at a fraction of the cost of whole genome sequencing. As of mid-2026, the exome and targeted sequencing landscape has matured beyond simple coding-region capture to encompass extended exome designs, custom panel engineering, and integrated methylation analysis — offering a graduated set of tools matched to specific research and clinical questions.

Exome and Targeted Sequencing Decision TreeFigure 1: The Exome and Targeted Sequencing Decision Tree — From Broad Discovery to Focused Validation

The Exome Advantage — 1% of the Genome, 85% of Disease Variants

The rationale for exome sequencing rests on a remarkably consistent observation across population-scale variant databases: the coding exome, despite its small genomic footprint, is disproportionately enriched for high-effect variants. In the gnomAD v4.1 dataset encompassing over 800,000 individuals, approximately 85 percent of variants classified as pathogenic or likely pathogenic in ClinVar fall within coding exons or canonical splice sites. The remaining 15 percent are distributed across promoters, enhancers, deep intronic regions, and non-coding RNAs — regions that are invisible to standard exome capture but accessible by whole genome sequencing or extended exome designs.

This enrichment is not an artifact of ascertainment bias — although clinical sequencing has historically targeted the exome — but reflects the fundamental biology that altering a protein's amino acid sequence generally produces a larger phenotypic effect than altering a regulatory element. A missense variant in a sodium channel gene can cause a cardiac arrhythmia; a synonymous variant in a deep intronic enhancer typically does not. The practical consequence is that exome sequencing, at approximately $220 to $260 per sample at academic core facilities as of mid-2026, delivers roughly 85 percent of the clinically actionable variant yield of whole genome sequencing at 30 to 50 percent of the cost.

The cost difference is driven by sequencing volume. A standard 30× to 50× exome generates roughly 5 to 10 gigabases of sequencing data per sample, compared to approximately 90 to 100 gigabases for a 30× whole genome. The smaller data footprint also reduces the computational burden for storage, alignment, and variant calling — an increasingly important consideration as laboratories scale from hundreds to thousands of samples.

Human WES ApplicationsFigure 2: Human WES Applications — Rare Disease Diagnosis, Cancer Genomics, and Population Genetics

Human Whole Exome Sequencing — Rare Disease, Cancer, and Population Genetics

Rare Disease Diagnosis and Trio Analysis

The highest-impact application of whole exome sequencing remains the diagnosis of rare Mendelian disorders. A 2025 randomized controlled trial of 653 trio families (Ungar et al., Genetics in Medicine) demonstrated that trio exome sequencing achieved a diagnostic yield of 35.9 percent, statistically indistinguishable from trio genome sequencing at 32.7 percent in the same population — while costing approximately 50 percent less (CAD $2,889 vs. $4,364 per trio). A parallel systematic review by Pandey et al. (2025) aggregating 108 studies and 24,631 pediatric probands reported a pooled diagnostic yield of 34.2 percent for exome or genome sequencing, compared to 18.1 percent for chromosomal microarray and targeted panel approaches.

Trio design — sequencing the affected proband together with both biological parents — is the single most impactful methodological choice in rare disease exome sequencing. By enabling immediate identification of de novo variants, phasing of compound heterozygous mutations, and rapid filtering of benign inherited variants, trio analysis adds approximately 15 to 20 percentage points to the diagnostic yield compared to proband-only sequencing. A 2025 analysis of 1,000 clinical trio cases (Malmgren et al., Frontiers in Genetics) reported an overall diagnostic rate of 39 percent, with de novo dominant variants accounting for 46 percent of solved cases and syndromic neurodevelopmental disorders achieving yields of 46 percent. When previously inconclusive proband-only cases were reanalyzed with parental samples, an additional 30 percent were solved.

A 2025 study of 137 Indian children with autism spectrum disorder (Bajaj et al., Journal of Human Genetics) reported a definitive diagnostic yield of 16.1 percent by trio exome sequencing, rising to 35 percent in syndromic ASD, with clinical management changes in 27.3 percent of diagnosed cases. A 2026 study of pediatric muscular disorders reported that exome sequencing as first-tier testing achieved a 72.3 percent diagnostic yield and eliminated the need for muscle biopsy in all cases during the seven-year study period.

Cancer Genomics — Tumor-Normal Pairs and Somatic Variant Detection

In cancer genomics, exome sequencing of matched tumor-normal sample pairs enables systematic identification of somatic single-nucleotide variants, small insertions and deletions, and copy-number alterations across the coding genome. The tumor-normal design — sequencing DNA from both the tumor specimen and a matched normal sample — distinguishes somatic mutations acquired during oncogenesis from germline variants present in all cells.

The analytical sensitivity of tumor-normal exome sequencing depends on sequencing depth and tumor purity. At a standard depth of 100× to 150× for the tumor sample and 50× for the normal sample, somatic variants present at allele fractions above 10 percent are reliably detected. For subclonal mutations at 5 to 10 percent variant allele frequency, deeper sequencing of 200× to 300× is recommended, and for circulating tumor DNA applications with variant fractions below 1 percent, targeted region sequencing with depths of 500× to 1,000× is required. A 2025 health economics analysis of genomic testing in oncology found that the per-sample incremental cost of exome sequencing over multi-gene panels was offset by the broader variant detection scope, particularly when results informed therapy selection or clinical trial eligibility.

Population Genetics and Biobank-Scale Studies

At population scale, exome sequencing enables cost-efficient characterization of coding variation across large cohorts. The UK Biobank exome sequencing effort, which released exome data for 470,000 participants in 2024-2025, has identified over 10 million coding variants and enabled gene-based rare variant association studies for thousands of traits. The All of Us Research Program and the NIH CARDIA study have adopted similar exome-first strategies for diverse population cohorts. At a per-sample cost of approximately $220 for library preparation and sequencing, exome sequencing allows cohorts of tens to hundreds of thousands of participants to be profiled within budgets that would be prohibitive for whole genome approaches.

Non-Human Exome SequencingFigure 3: Non-Human Exome Sequencing — Custom Bait Design for Animal and Plant Species

Animal and Plant Exome Sequencing — Custom Bait Design for Non-Human Species

While the human exome benefits from decades of genome annotation and commercially optimized capture kits, exome sequencing in non-human species requires custom bait design — the computational selection and synthesis of oligonucleotide probes targeting the coding regions of a specific organism's genome. This capability has opened exome-level variant discovery to agricultural genetics, evolutionary biology, and veterinary research.

The custom bait design process begins with the target organism's reference genome assembly and gene annotations. Coding exon coordinates are extracted, and overlapping probes of 80 to 120 base pairs are designed with optimized tiling density, GC content balancing, and repeat masking. For species with well-annotated genomes — mouse, rat, zebrafish, rice, maize, soybean — probe design can leverage Ensembl or NCBI gene models. For non-model organisms with draft genomes, RNA-seq data can supplement or substitute for gene annotation by identifying transcribed regions for probe targeting.

The economics differ from human exome sequencing. Custom bait panels incur an upfront synthesis cost — typically $2,000 to $10,000 depending on target size and probe count — that must be amortized across the sample cohort. For studies of fewer than 50 samples, the per-sample cost including bait synthesis can exceed that of low-coverage whole genome sequencing. For studies exceeding 100 samples with the same bait panel, the per-sample cost approaches the human exome benchmark. Oligo pool synthesis platforms as of 2026 can generate up to 4.35 million unique probe sequences per synthesis run, enabling exome-scale capture designs for virtually any eukaryotic genome.

Applications in animal genetics include mapping of domestication genes and production traits in livestock, identification of disease-causing variants in companion animals and veterinary species, and population genomic studies of wild and endangered species where non-invasive sampling limits DNA quantity. In plant genetics, exome capture has been used to characterize natural variation in crop wild relatives, identify alleles underlying agronomic traits in breeding populations, and perform variant discovery in polyploid genomes such as wheat and strawberry where whole genome sequencing remains computationally challenging. For projects targeting non-human species, custom gene panel design can complement exome-scale approaches by focusing on trait-specific gene sets once candidate regions are identified.

Custom Targeted Panel DesignFigure 4: Custom Targeted Panel Design — From Gene Selection to Validated Clinical Assay

Custom Targeted Panels — Designing for Specific Gene Sets and Genomic Regions

When the research question narrows to a specific gene set — a cancer predisposition panel of 50 to 100 genes, a cardiomyopathy panel of 30 to 50 genes, a pharmacogenomics panel of 10 to 20 pharmacokinetic genes — targeted panel sequencing offers higher depth, lower cost, and simpler analysis than exome sequencing. Custom panel design, in which the researcher specifies the exact genomic regions to be captured rather than relying on an off-the-shelf commercial kit, provides maximal flexibility for specialized applications.

The core design decision is the enrichment chemistry: hybrid capture or amplicon-based targeting. Hybrid capture uses biotinylated oligonucleotide probes to pull down complementary genomic fragments from a sequencing library, typically achieving on-target rates of 80 to 95 percent and supporting panels from a few dozen genes to several thousand. It tolerates DNA degradation well — important for formalin-fixed paraffin-embedded (FFPE) specimens — and captures structural variants and gene fusions when probes span intronic breakpoint regions. Amplicon-based methods use multiplex PCR to amplify target regions directly from genomic DNA, offering faster turnaround and lower input requirements (1 to 10 nanograms) but with reduced tolerance for primer-binding site variants and generally lower uniformity in high-GC regions. For clinical oncology panels where detection of gene fusions is critical, hybrid capture is the predominant choice.

Probe design quality distinguishes successful custom panels from failed ones. Key parameters as of 2026 include AI-assisted GC content optimization — probe sequences with GC content below 25 percent or above 75 percent are flagged for compensatory density increases or alternate-strand redesign — and pseudogene cross-reactivity screening, in which candidate probes are BLASTed against the reference genome to identify those with high sequence identity to processed pseudogenes. For genes with known pseudogene paralogs such as PMS2, SMN1/SM2, and CYP2D6, probes must be positioned to discriminate the functional gene from its pseudogene copies.

A 2025 real-world validation of a customized 69-gene hybrid capture RNA sequencing panel for B-lineage acute lymphoblastic leukemia (Sreedharanunni et al., Molecular Diagnosis and Therapy) achieved a 43 percent fusion detection rate in cases missed by standard FISH panels, demonstrating the clinical value of custom panel design for molecular subtyping. The emerging standard for clinical-grade custom panels specifies on-target rates above 90 percent, coverage uniformity with at least 95 percent of target bases exceeding 20 percent of mean depth, and analytical sensitivity for variants at 5 percent allele frequency at 500× mean depth. For researchers developing custom panels, bioinformatics services for probe design and in silico validation provide quality assurance before synthesis and wet-lab testing.

Targeted Bisulfite SequencingFigure 5: Targeted Bisulfite Sequencing and the Integrated Exome-to-Epigenome Workflow

Targeted Methylation Analysis — Single-Base Resolution at Specific Genomic Loci

Whole genome bisulfite sequencing remains prohibitively expensive for most studies requiring high-depth methylation analysis at specific loci. Targeted bisulfite sequencing — combining bisulfite conversion of unmethylated cytosines with capture or amplification of defined genomic regions — provides single-base resolution methylation data at a fraction of the cost, making it accessible for biomarker validation, epigenetic epidemiology, and functional studies of gene regulation.

The workflow begins with bisulfite treatment of genomic DNA, which converts unmethylated cytosines to uracils while leaving methylated cytosines unchanged. Following conversion, target regions are enriched either by hybrid capture probes designed against the bisulfite-converted reference sequence or by bisulfite-specific PCR primers that amplify both methylated and unmethylated alleles. Sequencing to depths of 500× to 1,000× enables quantification of methylation percentages at individual CpG sites with confidence intervals of approximately 2 to 5 percent.

A 2025 study demonstrated the cost-effectiveness of a targeted long-read bisulfite sequencing approach on the Oxford Nanopore platform (Pereyra et al., BMC Medical Genomics), amplifying over 1 kilobase fragments from 12 gene promoters to achieve single-base resolution across 788 CpG sites at a mean depth of 619× per site. The long-read format enables phasing of methylation patterns across adjacent CpG sites — information lost in short-read bisulfite sequencing — which is relevant for understanding coordinated epigenetic silencing at promoter CpG islands. A parallel 2025 study validated a multiplex PCR-based targeted methylation panel covering 141 CpG sites across 20 genes for regulatory T cell characterization, demonstrating high concordance with methylation-sensitive high-resolution melting analysis.

For researchers integrating genetic and epigenetic analysis, the combination of exome sequencing for variant discovery with targeted bisulfite sequencing for methylation validation at candidate loci provides a cost-efficient path from genome-wide screening to mechanistic follow-up. This integrated workflow is particularly relevant for imprinting disorders, where the pathogenic variant and its parent-of-origin-specific methylation pattern must both be characterized, and for cancer studies investigating the interplay between somatic mutations and promoter hypermethylation at tumor suppressor genes.

Service Workflow — From Sample to Biological Insight

A typical exome or targeted sequencing project proceeds through four stages. In the first stage, sample preparation, genomic DNA is extracted from the source material — blood, saliva, fresh-frozen tissue, FFPE sections, or extracted DNA provided by the researcher — and assessed for concentration, purity, and integrity. For FFPE samples, the extent of DNA crosslinking and fragmentation determines the feasibility of exome-scale capture versus targeted panel approaches. The DV200 metric, measuring the percentage of DNA fragments above 200 base pairs, has become the standard QC parameter for FFPE-derived DNA.

The second stage, library preparation and target enrichment, encompasses DNA fragmentation to 150 to 300 base pair inserts, end repair, A-tailing, adapter ligation, and PCR amplification, followed by hybridization to the capture probes — whether a commercial human exome kit, a custom bait panel for a non-human species, or a targeted gene panel. Post-capture amplification and pooling of multiplexed samples complete the library preparation workflow. Current exome capture kits from leading manufacturers (Twist Bioscience, Agilent, IDT, Roche) support 8-plex to 12-plex sample pooling per capture reaction, reducing per-sample library preparation costs by 30 to 40 percent compared to single-plex capture.

The third stage, sequencing, is typically performed on Illumina NovaSeq X Plus or equivalent platforms generating 2 × 150 base pair paired-end reads. The sequencing depth is tailored to the application: 50× to 100× mean target coverage for germline variant detection in constitutional DNA, 100× to 150× for tumor samples in somatic variant detection, 500× to 1,000× for targeted panels requiring low-frequency variant detection, and 500× to 1,000× for targeted bisulfite sequencing requiring precise methylation quantification.

The fourth stage, bioinformatic analysis, includes read alignment to the appropriate reference genome, duplicate marking or removal, base quality score recalibration, and variant calling using GATK HaplotypeCaller for germline variants or MuTect2 and VarScan2 for somatic variants. Variant annotation integrates population frequency databases (gnomAD, 1000 Genomes), clinical variant databases (ClinVar, HGMD), and in silico effect prediction tools (CADD, REVEL, SpliceAI). For trio analyses, automated de novo calling and compound heterozygote phasing are standard pipeline components. For whole exome sequencing data analysis, automated tiered variant filtration pipelines reduce the candidate variant list from tens of thousands to dozens, enabling efficient manual curation by clinical analysts.

FAQ

What is the difference between whole exome sequencing and whole genome sequencing?

Whole exome sequencing captures and sequences the approximately 1 percent of the genome comprising protein-coding exons, while whole genome sequencing reads the entire genome including intergenic regions, introns, and regulatory elements. Exome sequencing costs approximately $220 to $260 per sample (academic pricing, 2026) versus $400 to $700 for a 30× genome, and generates 5 to 10 gigabases of data versus 90 to 100 gigabases. The exome captures roughly 85 percent of known disease-causing variants, making it cost-effective for most diagnostic and research applications.

How many samples can be pooled in one exome capture reaction?

Current commercial exome capture kits support 8-plex to 12-plex sample pooling per capture reaction. Pooling 8 samples per capture reduces the per-sample library preparation cost by 30 to 40 percent compared to single-plex capture. The optimal multiplexing factor balances cost savings against the risk of sample-to-sample cross-contamination during hybridization.

What sequencing depth is needed for exome sequencing?

For germline variant detection in constitutional DNA, 50× to 100× mean target coverage is standard. For tumor-normal pairs in cancer genomics, 100× to 150× for the tumor and 50× for the normal sample is typical. For targeted panels requiring detection of variants below 5 percent allele frequency, 500× to 1,000× depth is recommended.

What is a trio exome and why does it improve diagnostic yield?

A trio exome sequences the affected individual (proband) together with both biological parents. This design enables immediate identification of de novo variants — those present in the proband but absent from both parents — and confirmation of compound heterozygous variants where each parent contributes one mutated allele. Trio analysis adds approximately 15 to 20 percentage points to diagnostic yield compared to proband-only exome sequencing.

Can exome sequencing be performed on non-human species?

Yes. Custom bait panels can be designed for any species with an available reference genome and gene annotation. Probe design tools extract coding exon coordinates and design capture oligonucleotides specific to the target organism. Custom bait panels for non-model organisms incur an upfront synthesis cost that must be amortized across the sample cohort, but for studies of 100 or more samples, the per-sample cost approaches that of human exome sequencing.

What is the difference between hybrid capture and amplicon-based targeted sequencing?

Hybrid capture uses biotinylated oligonucleotide probes to enrich target regions from a sequencing library, achieving high on-target rates (80 to 95 percent), supporting large panels (dozens to thousands of genes), and enabling detection of structural variants and gene fusions. Amplicon-based methods use multiplex PCR for direct target amplification, offering faster turnaround and lower input requirements but reduced tolerance for primer-binding site variants and typically lower uniformity. Hybrid capture is preferred for clinical oncology panels; amplicon panels are effective for rapid hotspot screening.

When should I choose a custom targeted panel over whole exome sequencing?

Choose a custom panel when the research or clinical question is focused on a defined gene set — for example, a 50-gene cardiomyopathy panel, a 100-gene cancer predisposition panel, or a 20-gene pharmacogenomics panel. Targeted panels offer higher depth (500× to 1,000× vs. 50× to 100× for exomes) at lower cost per sample, simpler analysis with fewer incidental findings, and faster turnaround. Choose exome sequencing when the gene set is broad or undefined, when discovery of novel gene-disease associations is a goal, or when the differential diagnosis spans multiple organ systems.

What is targeted bisulfite sequencing?

Targeted bisulfite sequencing combines bisulfite conversion of unmethylated cytosines with capture or amplification of specific genomic regions to quantify DNA methylation at single-base resolution. It provides the methylation information of whole genome bisulfite sequencing at a fraction of the cost by focusing on predefined loci — typically promoter CpG islands, differentially methylated regions, or imprinting control regions — with sequencing depths of 500× to 1,000×.

References:

  1. Ungar WJ, Wu V, Marshall CR, et al. A microcosting and cost consequence analysis from a randomized controlled trial comparing genome sequencing with exome sequencing for genetic diagnosis. Genetics in Medicine. 2026;28(2):101561. https://doi.org/10.1016/j.gim.2025.101561
  2. Kaschta D, Post C, Gaass F, et al. Evaluating genome sequencing strategies: trio, singleton, and standard testing in rare disease diagnosis. Genome Medicine. 2025;17:100. https://doi.org/10.1186/s13073-025-01516-7
  3. Bajaj S, Gandhi S, Geetha TS, et al. Prospective study to analyze the yield and clinical impact of trio exome sequencing in 137 Indian children with autism spectrum disorder. Journal of Human Genetics. 2025;70:611-624. https://doi.org/10.1038/s10038-025-01368-4
  4. Malmgren H, Kvarnung M, Gustafsson P, et al. Diagnostic yield of 1000 trio analyses with exome and genome sequencing in a clinical setting. Frontiers in Genetics. 2025;16:1580879. https://doi.org/10.3389/fgene.2025.1580879
  5. Miya F, Nakato D, Suzuki H, et al. Augmenting cost-effectiveness in clinical diagnosis using extended whole-exome sequencing: SNVs, SVs, and beyond. Journal of Human Genetics. 2026;71:13-21. https://doi.org/10.1038/s10038-025-01403-4
  6. Sreedharanunni S, Thakur V, Balakrishnan A, et al. Effective utilization of a customized targeted hybrid capture RNA sequencing in the routine molecular categorization of adolescent and adult B-lineage acute lymphoblastic leukemia. Molecular Diagnosis and Therapy. 2025;29(3):407-418. https://doi.org/10.1007/s40291-025-00779-5
  7. Pereyra S, Sardina A, Neumann R, et al. Cost-effective promoter methylation analysis via target long-read bisulfite sequencing: a case study in severe preterm birth. BMC Medical Genomics. 2025;18:122. https://doi.org/10.1186/s12920-025-02193-6

For research purposes only, not intended for clinical diagnosis, treatment, or individual health assessments.

For research purposes only, not intended for clinical diagnosis, treatment, or individual health assessments.
Speak to Our Scientists
What would you like to discuss?
With whom will we be speaking?

* is a required item.

Contact CD Genomics
Terms & Conditions | Privacy Policy | Feedback   Copyright © CD Genomics. All rights reserved.
Top