Animal and Plant Exome Sequencing: Custom Target Capture for Non-Human Species
Whole genome sequencing at the population scale remains cost-prohibitive for many agricultural and ecological research programs. A single 30× cattle genome generates roughly 80 to 90 gigabases of sequencing data at a reagent cost of approximately $400 to $600 per sample. For a study of 500 animals — a modest cohort by modern breeding-program standards — the sequencing budget alone approaches $250,000. When the genome in question is hexaploid wheat at 14.6 gigabases or conifer at 20 gigabases, the economics become untenable. Whole exome sequencing — capturing and sequencing only the protein-coding fraction of the genome, typically 1 to 2 percent of total genome size — offers a practical alternative that preserves the vast majority of functionally interpretable variation at a fraction of the cost. This article examines the custom bait design, cross-species capture considerations, bioinformatic requirements, and cost-effectiveness of exome sequencing applied to non-human species, from livestock and crops to wild populations and model organisms.
Figure 1: The Non-Human Exome Sequencing Decision Framework — When to Choose Exome Over Genome
You may be interested in
Learn more
When Whole Genome Sequencing Is Too Much — The Exome as a Practical Alternative
The rationale for exome sequencing in non-human species is fundamentally economic, but it is reinforced by biology. Across vertebrates, the protein-coding exome accounts for roughly 1 to 2 percent of the genome yet harbors the vast majority of high-effect variants underlying phenotypic traits. In cattle, the annotated coding sequence spans approximately 33 to 35 megabases out of a 2.7 gigabase genome. In hexaploid wheat, roughly 132 megabases of coding sequence are distributed across a 14.6 gigabase genome that is approximately 85 percent repetitive. Sequencing the entire genome to sufficient depth for rare variant detection in these species demands a sequencing investment that scales with genome size rather than with functional content. Exome sequencing decouples the two: it delivers 50× to 100× coverage of coding regions regardless of genome size, producing 3 to 10 gigabases of data per sample rather than 80 to 300 gigabases.
This decoupling is particularly valuable for species with large or polyploid genomes where whole genome sequencing depth is constrained by budget. A 2025 review of ruminant livestock genomics (Xu et al., Science China Life Sciences) documented the rapid expansion of reference-quality genome assemblies and population-scale resequencing efforts in cattle, sheep, and goat, noting that while whole genome resequencing of large breeding cohorts remains expensive, focused genotyping and sequencing strategies have emerged as practical bridges between reference assembly and functional variant discovery. Similarly, plant breeding programs working with wheat, barley, oat, and conifer species have adopted exome capture as a primary variant discovery platform, generating population-scale coding variant catalogs that would be economically infeasible via whole genome sequencing.
The decision to use exome rather than genome sequencing turns on the research question. For studies focused on coding variants — domestication sweeps in livestock, disease resistance alleles in crops, nonsynonymous mutations underlying metabolic adaptations — exome sequencing captures the relevant variation at roughly 10 to 20 percent of the sequencing cost. For studies requiring structural variant detection, regulatory element characterization, or de novo genome assembly, whole genome sequencing remains the appropriate tool. A growing number of agricultural genomics programs adopt a hybrid strategy: deep exome sequencing (50× to 100×) on the full cohort for coding variant discovery, combined with low-coverage whole genome sequencing (1× to 5×) on a subset for structural variant imputation. This approach captures the majority of both coding and structural variation at roughly 40 to 60 percent of the cost of high-coverage whole genome sequencing on the full cohort.
Applications of non-human exome sequencing span domestication genetics (identifying selective sweeps associated with artificial selection in livestock and crops), disease resistance mapping (cloning genes conferring resistance to pathogens in agricultural species), metabolic pathway analysis (characterizing variation in cytochrome P450 and secondary metabolite genes in plants), conservation genomics (assessing genetic diversity in endangered populations from noninvasive samples), and phylogenomics (resolving evolutionary relationships using coding-region markers across divergent taxa). A 2025 study of wild chimpanzees across six African regions (Hayakawa et al., Scientific Reports) demonstrated the utility of human exome baits for population genomic analysis of noninvasive fecal samples, achieving on-target rates of approximately 80 percent despite the roughly 6 million years of divergence between humans and chimpanzees.
Figure 2: Custom Bait Design Workflow — From Reference Genome to Species-Specific Exome Capture Panel
Custom Bait Design for Non-Human Exomes — From Reference Genome to Probe Synthesis
The feasibility of non-human exome sequencing rests on custom bait design — the computational selection and chemical synthesis of oligonucleotide probes that selectively hybridize to the coding regions of a target species' genome. Unlike human exome sequencing, for which multiple commercial kits with decades of iterative optimization are available, non-human exome capture typically requires a custom probe panel designed specifically for the target organism or a closely related species.
The design workflow begins with the target species' reference genome assembly and gene annotations. Protein-coding exon coordinates are extracted from GFF3 or GTF annotation files, and overlapping probes of 80 to 120 base pairs are designed to tile across each target exon. Key design parameters include probe length — longer probes (100 to 120 base pairs) tolerate more sequence divergence and are preferred for cross-species applications — and tiling density, typically 1.5× to 2× coverage of each target base to ensure uniform capture even if individual probes fail. GC content is balanced across the probe set, with extremes below 25 percent or above 75 percent flagged for compensatory density increases or strand-switching redesign, as extreme GC content reduces hybridization efficiency and contributes to coverage gaps.
Repeat masking is a critical quality-control step in non-human exome probe design. Many crop genomes — wheat, maize, soybean — contain extensive repetitive elements that, if included in the probe set, cause off-target capture and reduce effective sequencing depth on coding regions. Candidate probes are screened against the reference genome using BLAST or similar alignment tools, and those with significant secondary alignments to non-target regions are excluded. For genes with processed pseudogenes — a particular concern in polyploid crop genomes where subgenome duplications create near-identical paralogs — probes must be positioned to discriminate the functional gene copy from its pseudogene and homeolog counterparts.
The choice of probe synthesis platform affects cross-species capture performance. The two dominant chemistries — solution-based hybrid capture with biotinylated RNA or DNA probes — are available from multiple commercial suppliers. NimbleGen SeqCap probes (Roche) have been reported to outperform IDT xGen probes for cross-species applications at moderate phylogenetic distances. Webster et al. (2018, bioRxiv) demonstrated that NimbleGen baits captured over 90 percent of annotated coding sequence in lemurs — primates that diverged from humans approximately 60 million years ago — while IDT probes yielded substantially less callable sequence at equivalent sequencing depth.
For species without a high-quality reference genome, transcriptome-based probe design offers an alternative. RNA-seq data are assembled into a reference transcriptome, coding regions are predicted, and probes are designed against the predicted exons. This approach has been successfully applied to non-model amphibians, where FrogCap — a modular probe set for phylogenomics across all frogs (Hutter et al., Molecular Ecology Resources, 2022) — demonstrated that consensus-based probe design, in which probes are designed against a consensus sequence derived from multiple species, outperforms exemplar-based design (probes designed against a single reference genome) for cross-species capture sensitivity and specificity. The FrogCap framework allows researchers to select marker subsets appropriate to their phylogenetic scale of interest, from deep ordinal-level relationships to shallow population-level structure. For researchers developing custom capture panels, bioinformatics services for probe design and in silico validation provide critical quality assurance before synthesis and wet-lab testing.
The economics of custom bait design differ markedly from off-the-shelf human exome kits. Custom probe panel synthesis costs range from approximately $2,000 to $10,000 depending on target territory size and probe count, with oligo pool synthesis platforms as of 2026 capable of generating up to 4.35 million unique probe sequences per synthesis run. This upfront cost must be amortized across the sample cohort. For studies of fewer than 50 samples, the per-sample cost including amortized bait synthesis can exceed that of low-coverage whole genome sequencing. For studies exceeding 100 samples using the same bait panel, the per-sample cost approaches human exome benchmarks of approximately $200 to $300. This economic structure makes exome capture particularly attractive for breeding programs, biobank-scale cohorts, and long-term population monitoring studies where the initial bait investment is distributed across hundreds to thousands of samples.
Figure 3: Bioinformatics Pipeline for Non-Model Organism Exome Analysis — From Raw Reads to Annotated Variants
Bioinformatics for Non-Human Exomes — Adapting Standard Pipelines to Non-Model Genomes
The bioinformatic analysis of non-human exome data largely follows the same workflow as human exome analysis — read alignment, duplicate marking, base quality score recalibration, variant calling, and annotation — but with several adaptations specific to non-model organisms. The most consequential decision is the choice of reference genome for read alignment. When a high-quality reference assembly is available for the target species, reads are aligned to that assembly using BWA-MEM with standard parameters. When the target species lacks a reference genome, reads can be aligned to the nearest available relative's genome, but this introduces reference bias: reads from genomic regions that have diverged between the target and reference species map with lower efficiency, producing false-negative variant calls in the most biologically interesting — because most diverged — regions of the genome.
The pioneering work of Vallender (2011, Genome Biology) established the feasibility of cross-species exome alignment by demonstrating that human exome capture probes applied to chimpanzee and rhesus macaque samples produced coverage distributions comparable to human samples in coding regions, though untranslated region coverage declined with increasing phylogenetic distance. At approximately 25 million years of divergence — the human-macaque split — coding sequence coverage remained adequate for variant discovery, but UTR and flanking intronic coverage dropped substantially.
For variant calling in non-model organisms, GATK HaplotypeCaller in gVCF mode remains the standard, but hard filtering thresholds must be calibrated to the specific dataset because the truth-set training data required for Variant Quality Score Recalibration (VQSR) are typically unavailable for non-human species. Recommended hard-filter parameters for non-model exome data include QD below 2.0, Fisher Strand bias above 60.0, and mapping quality below 40.0. Variant calling in polyploid species adds further complexity, requiring tools capable of distinguishing homeolog-specific variants from sequencing errors — a challenge particularly acute in hexaploid wheat and tetraploid crops.
Variant annotation presents a parallel challenge: the standard human annotation databases (ClinVar, gnomAD, dbNSFP) are irrelevant for non-human species, and researchers must instead rely on SnpEff with a custom database built from the target species' genome annotation, supplemented by InterProScan for protein domain annotation and BLAST-based ortholog inference for cross-species functional prediction. The impact of phylogenetic distance on capture efficiency was systematically quantified by Lemarcis et al. (2025, Molecular Ecology Resources), who designed 1,125 exon capture probes for Neogastropoda (marine snails) and tested them across 150 specimens from 30 species spanning a broad range of divergence times. They reported a negative linear correlation between genetic distance (p-distance) and the number of exons successfully captured, but — contrary to the commonly cited 10 percent divergence threshold for capture failure — found no sharp cutoff. Instead, capture efficiency declined gradually, and the authors recommended augmenting probe sets with transcriptome-derived probes from divergent lineages rather than relying on a single reference-designed panel. This finding has practical implications for exome studies of species without a reference genome: while human or model-organism baits can capture functionally informative sequence from relatively distant relatives, the yield declines predictably with divergence, and probe sets should be iteratively refined as taxon-specific genomic resources become available.
Figure 4: Case Examples — Exome Sequencing Applications in Livestock and Crop Science
Case Examples — From Livestock Disease Resistance to Crop Improvement
Livestock Disease Resistance
Identifying genetic variants underlying disease resistance in livestock is a central goal of agricultural genomics, with direct economic and animal-welfare implications. Exome sequencing has been deployed to characterize coding variation in immune-related genes across cattle breeds, enabling the discovery of alleles associated with resistance to bovine respiratory disease, trypanosomiasis, and mastitis. The 2025 ruminant genomics review by Xu et al. (Science China Life Sciences) surveyed the accelerating progress in livestock reference genome assembly, population genomics, and multiomics integration that has enabled systematic discovery of functional variants underlying production traits, disease resistance, and environmental adaptation across cattle, sheep, and goat breeds.
The cost advantage of exome sequencing is particularly salient for livestock species with large population sizes. A dairy breeding program genotyping 2,000 bulls for genomic selection might budget $30,000 to $50,000 for exome capture and sequencing, compared to $150,000 to $300,000 for whole genome sequencing at comparable depth. When the research goal is cataloguing coding variants in immune genes, production trait loci, or known disease-associated pathways, the exome-first strategy preserves the vast majority of actionable information while respecting typical agricultural research budgets. For projects targeting specific gene sets in livestock, custom gene panel sequencing provides a complementary approach that focuses sequencing resources on validated trait-associated loci.
Crop Improvement — Wheat Disease Resistance
In crop genomics, exome capture has been transformative for species with large, polyploid genomes. Bread wheat (Triticum aestivum, 2n = 6x = 42) possesses a 14.6 gigabase genome in which approximately 132 megabases — less than 1 percent — represent coding sequence. Whole genome sequencing of wheat breeding populations remains prohibitively expensive, but exome capture panels covering the annotated wheat coding sequence enable cost-effective variant discovery at population scale.
A 2025 study published in Nature Communications (Yang et al.) demonstrated the power of this approach for cloning an agriculturally critical disease-resistance gene. The researchers performed GWAS and whole-exome capture sequencing on two nested bi-parental wheat populations segregating for resistance to Fusarium crown rot, a soil-borne fungal disease that reduces grain yield and contaminates harvests with mycotoxins. By combining exome-wide association mapping with transcriptomic profiling, they identified TaCAT2, encoding a catalase enzyme that detoxifies reactive oxygen species produced during fungal infection. The resistance haplotype TaCAT2-Ser214, distinguished by a single phosphorylation-site polymorphism, confers enhanced protein stability and superior hydrogen peroxide scavenging capacity without agronomic penalty. This gene was cloned from exome-capture data at a fraction of the cost that would have been required for whole genome sequencing of the same populations — roughly $8,000 for exome capture and sequencing of the mapping population, compared to an estimated $120,000 for equivalent-depth whole genome sequencing.
Beyond single-gene cloning, wheat exome capture has enabled population-scale surveys of coding variation. Commercial wheat exome panels covering over 100,000 genes with more than 2.5 million probes are now available, supporting GWAS for grain yield, quality, stress tolerance, and disease resistance across diverse wheat germplasm collections. Similar exome capture resources have been developed for barley, maize, soybean, rice, and cotton, with probe counts and target territories scaled to each species' genome size and annotation quality.
Conservation Genomics — Noninvasive Sampling in Wild Populations
A distinctive advantage of exome capture in conservation genomics is its compatibility with low-quality, low-quantity DNA from noninvasive samples. The 2025 chimpanzee population genomics study by Hayakawa et al. used fecal DNA collected from 42 wild chimpanzees across six African field sites, extracting DNA stored in lysis buffer at ambient temperature and capturing it with commercial human exome baits. Despite the degraded starting material and the human-chimpanzee divergence of approximately 6 million years, on-target rates saturated at approximately 80 percent. The exome-wide sequence data successfully discriminated local populations — resolving finer population structure than mitochondrial genomes alone — and identified candidate loci under selection related to pathogen pressure and dietary adaptation. The study reported that high-quality blood-derived samples required approximately 5 gigabases of sequencing to reach coverage saturation, while low-quality fecal samples required 10 to over 20 gigabases — a two- to fourfold penalty that nevertheless kept per-sample sequencing costs within the range of $100 to $200.
This compatibility with noninvasive sampling is relevant for studies of endangered species where blood or tissue collection is ethically or logistically infeasible. Feces, hair, shed feathers, and museum specimens — all amenable to exome capture with appropriate probe design — can provide population genomic data without handling or disturbing study subjects. The Webster et al. (2018) lemur study similarly demonstrated that human baits can generate high-quality exome data from wild primates sampled noninvasively, capturing over 90 percent of coding sequence at over 7× depth in species that last shared a common ancestor with humans approximately 60 million years ago.
Figure 5: Cost Comparison — Whole Genome Sequencing vs. Whole Exome Sequencing for Non-Human Species
Cost Comparison — WGS vs. Exome for Non-Human Species
The cost calculus for non-human exome sequencing differs from the human case in one critical respect: the upfront investment in custom bait design. For a researcher studying a species for which no commercial exome kit exists, the total project cost is the sum of bait synthesis (amortized across samples) plus per-sample library preparation and sequencing. The breakeven point — the sample number at which exome sequencing becomes cheaper than low-coverage whole genome sequencing for equivalent coding-region variant yield — depends on genome size, target territory size, and the researcher's tolerance for missing non-coding variants.
Representative costs for mid-2026 academic core facility pricing provide a framework. A 30× cattle whole genome generates approximately 80 gigabases of data at a library-plus-sequencing cost of roughly $400 to $500 per sample. A cattle exome (approximately 35 megabases target territory) captured at 80× mean depth generates roughly 5 to 7 gigabases at a library-plus-sequencing cost of roughly $150 to $200 per sample, plus the amortized bait synthesis cost. At a bait panel cost of $5,000 amortized across 100 samples ($50 per sample), the per-sample exome cost is approximately $200 to $250 — roughly half the cost of whole genome sequencing. At 500 samples ($10 per sample amortized), the exome cost drops to approximately $160 to $210 per sample.
For large-genome species, the exome advantage widens. A 30× wheat whole genome at 14.6 gigabases generates approximately 440 gigabases at a cost exceeding $1,500 per sample. A wheat exome (approximately 132 megabases target territory) captured at 80× generates roughly 15 to 20 gigabases at a cost of approximately $250 to $350 per sample including amortized bait synthesis — a four- to sixfold cost reduction. For conifer species with genomes exceeding 20 gigabases, the differential is larger still. These economics explain why wheat and conifer genomics communities have been early and enthusiastic adopters of exome capture for population-scale studies.
The trade-off is completeness. Exome sequencing does not detect regulatory variants in promoters and enhancers, structural variants larger than can be inferred from read-pair and split-read signals at exon boundaries, or variants in unannotated genes and non-coding RNAs. For breeding programs focused on coding variation in annotated genes, this is an acceptable limitation. For discovery-focused research programs interested in the full spectrum of genomic variation, whole genome sequencing — potentially at lower depth supplemented by imputation — remains the appropriate choice. The trend in agricultural and ecological genomics as of mid-2026 is toward stratified study designs: deep exome sequencing on the full cohort to capture coding variants with high sensitivity, supplemented by low-coverage whole genome sequencing on a representative subset for structural variant discovery and imputation. For researchers seeking to combine coding and structural variant detection economically, targeted region sequencing approaches that focus on specific genomic intervals can bridge the gap between exome-scale and genome-scale strategies.
FAQ
When should I choose exome sequencing over whole genome sequencing for non-human species?
Choose exome sequencing when your research question focuses on coding-region variants (nonsynonymous, splice-site, and frameshift mutations), when you are working with a species that has a large or polyploid genome making whole genome sequencing cost-prohibitive, or when you need to screen large populations (hundreds to thousands of individuals) for functional variants within a constrained budget. Choose whole genome sequencing when you need to detect structural variants, regulatory mutations, or variants in non-coding regions, or when performing de novo genome assembly.
Can human exome capture kits be used on non-human species?
Yes, within limits. Human exome baits have been successfully used on non-human primates including chimpanzees (~6 million years divergence, ~80 percent on-target), rhesus macaques (~25 million years, good CDS coverage but reduced UTR coverage), and even lemurs (~60 million years, over 90 percent of CDS captured using NimbleGen baits). Capture efficiency declines gradually with phylogenetic distance — there is no sharp cutoff — but probe sets optimized for the target species will always outperform cross-species applications of human kits. For species beyond primates, custom bait design is recommended.
How much does a custom exome bait panel cost?
Custom bait panel synthesis costs typically range from $2,000 to $10,000 depending on target territory size and probe count. This upfront cost is amortized across the sample cohort. For studies exceeding 100 samples, the per-sample bait cost drops below $50 to $100, making exome capture cost-competitive with whole genome sequencing for coding-region variant discovery.
What sequencing depth is needed for non-human exome sequencing?
For germline variant discovery in diploid species, 50× to 80× mean target coverage is standard — comparable to human exome recommendations. For polyploid species such as hexaploid wheat, higher depths of 80× to 120× are recommended to distinguish homeolog-specific variants. For low-frequency variant detection in pooled or heterogeneous samples, 150× to 300× depth is appropriate.
What bioinformatic tools are used for non-human exome analysis?
The core pipeline uses the same tools as human exome analysis: BWA-MEM for read alignment, GATK HaplotypeCaller for variant calling, and SnpEff for variant annotation. The critical difference is that SnpEff requires a custom database built from the target species' genome annotation, since pre-built databases are only available for major model organisms. Hard filtering thresholds must be empirically calibrated, as the truth-set training data required for VQSR are unavailable for most non-human species.
How does cross-species capture efficiency vary with phylogenetic distance?
Capture efficiency declines gradually with increasing genetic distance between the probe-design species and the target species, with no sharp threshold at which capture fails. A 2025 systematic study in marine snails (Lemarcis et al.) demonstrated a linear negative correlation between p-distance and exons captured. Probes designed from a consensus of multiple related species (consensus-based design) outperform those designed from a single reference genome (exemplar-based design) for cross-species applications.
Which species benefit most from exome sequencing over whole genome sequencing?
Species with large genomes (wheat at 14.6 Gb, conifers at 20+ Gb), polyploid genomes (hexaploid wheat, tetraploid crops), and large population study designs (hundreds to thousands of individuals) benefit most from the exome approach. The cost savings are proportional to genome size: roughly 2-fold for cattle (2.7 Gb), 4- to 6-fold for wheat, and 8-fold or more for conifers.
References:
- Vallender EJ. Expanding whole exome resequencing into non-human primates. Genome Biology. 2011;12(9):R87. https://doi.org/10.1186/gb-2011-12-9-r87
- Webster TH, Guevara EE, Lawler RR, Bradley BJ. Successful exome capture and sequencing in lemurs using human baits. bioRxiv. 2018. https://doi.org/10.1101/490839
- Hutter CR, Cobb KA, Portik DM, Travers SL, Wood PL Jr, Brown RM. FrogCap: A modular sequence capture probe-set for phylogenomics and population genetics for all frogs, assessed across multiple phylogenetic scales. Molecular Ecology Resources. 2022;22(3):1100-1119. https://doi.org/10.1111/1755-0998.13517
- Xu S, Akhatayeva Z, Liu GE, et al. Genetic advancements and future directions in ruminant livestock breeding: from reference genomes to multiomics innovations. Science China Life Sciences. 2025;68(4):934-960. https://doi.org/10.1007/s11427-024-2744-4
- Lemarcis T, Blin A, Cariou M, Derzelle A, Farhat S, Fedosov A, Zaharias P, Zuccon D, Puillandre N. Too Far From Relatives? Impact of the Genetic Distance on the Success of Exon Capture in Phylogenomics. Molecular Ecology Resources. 2025;25(4):e14064. https://doi.org/10.1111/1755-0998.14064
- Hayakawa T, Kishida T, Go Y, Inoue E, Kawaguchi E, Aizu T, Ishizaki H, Toyoda A, Fujiyama A, Matsuzawa T, Hashimoto C, Furuichi T, Agata K. Genome-scale evolution in local populations of wild chimpanzees. Scientific Reports. 2025;15:548. https://doi.org/10.1038/s41598-024-84163-z
- Yang X, Zhang L, Wei J, Liu L, Liu D, Yan X, Yuan M, Zhang L, Zhang N, Ren Y, Chen F. A TaSnRK1α-TaCAT2 model mediates resistance to Fusarium crown rot by scavenging ROS in common wheat. Nature Communications. 2025;16:2549. https://doi.org/10.1038/s41467-025-57936-x
Related Services
For research purposes only, not intended for clinical diagnosis, treatment, or individual health assessments.