When genome-wide discovery has narrowed the search to known genes, GWAS loci, QTL intervals, or custom regions, focused cohort sequencing turns those targets into variant and population-level evidence.
When genes, loci, or candidate intervals are already known, a targeted resequencing service concentrates sequencing on those predefined regions instead of distributing reads across the entire genome. This makes it practical to investigate common and rare variants, compare allele frequencies, and evaluate regional haplotypes across hundreds to thousands of samples while keeping the resulting dataset focused and manageable.
CD Genomics supports population-scale studies that begin with gene lists, genomic coordinates, BED files, GWAS loci, QTL candidate intervals, exons, or regulatory regions. Project design can cover target definition, enrichment strategy, sequencing, cohort-level quality control, variant analysis, and study-ready outputs. The final assay configuration and sequencing allocation are set according to target size, genome context, sample quality, cohort structure, and the biological question.
Planning a large-cohort follow-up study? Prepare your species, reference assembly, cohort structure, target list, and expected analyses for a project-design discussion.
Figure 1: Targeted resequencing connects a defined target list with focused sequencing and population-level interpretation across a large cohort.
Targeted resequencing is next-generation sequencing of genomic regions selected before library enrichment or amplification. Unlike WGS, which surveys the full genome, or WES, which is limited to annotated coding exons, a custom design can focus on the exact genes, exons, regulatory regions, candidate loci, or coordinates that answer the study question.
The value is not simply deeper sequencing. The important advantage is controlled scope: reads are concentrated on relevant targets, which reduces unnecessary data and supports coordinated analysis across many samples. In practice, this is useful after discovery work has already narrowed the search space and the next decision depends on validating variants, estimating population frequencies, or comparing the same regions across groups.
| Target Input | Typical Use | Research Benefit |
| Candidate genes or exons | Follow-up of prior functional evidence or phenotype-linked genes | Concentrates analysis on interpretable gene regions, reducing the number of unrelated variants that require review. |
| GWAS loci or fine-mapping intervals | Targeted sequencing after association discovery | Adds sequence-level variation around associated loci, helping researchers move from a signal to candidate variants and regional haplotypes. |
| QTL candidate intervals | Trait mapping follow-up in plant or animal populations | Tests the same candidate interval across a broader cohort, supporting cross-population validation and marker prioritization. |
| Regulatory or noncoding regions | Promoters, enhancers, conserved elements, and other defined features | Extends the study beyond coding sequence when the hypothesis depends on regulatory variation that WES would not capture. |
| Custom genomic coordinates | BED files or assembly-specific interval lists | Preserves the investigator's exact discovery boundaries, which is important when regions do not align with standard gene annotations. |
Hybridization capture uses complementary probes to enrich many dispersed or extended genomic regions from a sequencing library. Its flexible target architecture can accommodate genes, exons, and noncoding intervals in one design, so it is often a strong fit when the study spans numerous or irregular targets.
Amplicon-based targeting uses multiplex PCR to amplify defined regions. It can be efficient for compact, well-characterized targets and very large sample counts, so the method may be preferable when the total target set is small and primer design is straightforward. Repetitive sequence, GC extremes, nearby homologs, target length, and expected variant types all influence the choice; neither method is universally superior.
A large-cohort project succeeds when target design, sample processing, and analysis use the same coordinate system and QC logic from the beginning. The workflow below treats each sample as part of a cohort rather than as an isolated sequencing library, which helps prevent technical variation from being mistaken for population variation.
1. Study and target review
We review the biological objective, reference assembly, target list, population groups, sample count, and intended downstream analyses. This connects assay scope to the final decision the study must support, rather than treating every requested coordinate as equally informative.
2. Target-design feasibility assessment
Candidate regions are assessed for mappability, repeats, GC composition, paralogous sequence, interval boundaries, and compatibility with the selected targeting method. Identifying difficult regions before production reduces avoidable coverage gaps and shows where orthogonal validation or an alternative method may be needed.
3. Sample and library quality control
DNA identity, quantity, purity, and integrity are evaluated using project-appropriate checks before library preparation. Applying consistent acceptance logic across the cohort reduces batch imbalance and limits downstream missingness caused by uneven starting material.
4. Target enrichment and sequencing
Libraries are enriched by the agreed capture or amplicon strategy and sequenced with allocation matched to target size and analysis goals. Because allocation is based on the study design, coverage can be interpreted against the variants and cohort comparisons that matter rather than against a generic depth claim.
5. Data processing and cohort harmonization
Reads are checked, aligned to the confirmed reference, and evaluated for target coverage, uniformity, mapping behavior, and sample-level consistency. Variant processing then follows a shared cohort framework, which improves comparability among groups and makes exclusions traceable.
6. Review and delivery
Results are reviewed against target completeness, sample-level QC, and the agreed analysis scope before delivery. This links each file and figure to a documented method and QC record, making the dataset easier to use in downstream association, frequency, or haplotype analyses.
Figure 2: The cohort workflow links every technical step to a decision point and a documented output.
Useful targeted data require more than a high average depth. Sample-level callability, per-target coverage, on-target behavior, uniformity, missingness, and cross-batch consistency show whether a locus can be compared fairly across the cohort. These measurements help researchers identify samples or regions that need cautious interpretation before frequency or association analysis.
The published example below shows why coverage should be viewed across all targets and samples rather than summarized as one number.
Figure 3: Original evidence summary showing why target coverage should be reviewed across samples and loci, based on the independent study by Oh et al. (2023).
The most useful submission package combines biological samples with target and cohort metadata. Exact DNA requirements depend on species, genome complexity, target architecture, and the selected library strategy, so final acceptance criteria should be confirmed during project design rather than inferred from a generic threshold.
| Input | What to Provide | Review Point | Why It Matters |
| Genomic DNA | DNA from the study species, submitted in the agreed container and buffer | Quantity, purity, integrity, contamination risk, and sample consistency | Comparable starting material reduces uneven library performance across groups. |
| Source material | Tissue or another source type when extraction is included in scope | Species, preservation method, storage history, and extraction feasibility | Early review reveals degradation or inhibitor risks before cohort processing begins. |
| Target definition | Gene symbols, transcript IDs, genomic coordinates, BED file, or candidate intervals | Reference assembly version, strand/feature definition, interval padding, and duplicate regions | A confirmed coordinate system prevents probe design and variant annotation from referring to different loci. |
| Cohort metadata | Sample IDs, populations or groups, species or ancestry information, collection site, and relevant covariates | Balanced groups, missing fields, batch structure, and permitted comparisons | Well-structured metadata allows QC and population summaries to answer the planned biological comparisons. |
| Discovery evidence | Prior WGS, WES, GWAS, QTL, literature, or pilot results | How each target was nominated and which variants or regions require follow-up | Traceable target selection keeps the panel aligned with the original discovery question. |
Bioinformatics converts focused reads into cohort-ready evidence. The minimum workflow evaluates read quality, alignment, target performance, and variants; optional modules extend the dataset into allele-frequency, haplotype, population-comparison, or association-ready outputs when the study design supports them.
Representative visuals should make both technical quality and biological interpretation visible. A useful demo set includes a target-coverage heatmap, per-sample callability or missingness summary, allele-frequency comparison, regional variant view, haplotype-frequency plot, and an optional association or population-differentiation view. Final plot types depend on the agreed scope; illustrative figures must not be interpreted as guaranteed project results.
Figure 4: Illustrative analysis views show how target performance and population-level findings can be reviewed together.
Deliverables are scoped around reuse: raw data support reprocessing, processed files support variant review, and cohort summaries support biological interpretation. Providing these layers together means technical teams can audit the analysis while investigators can move directly to the population-level questions.
This approach is strongest when discovery has already narrowed the genomic search space and the next question requires more samples, more consistent coverage of nominated regions, or both. The common thread is a defined target set and a population-level decision.
Associated loci or mapped intervals can be resequenced across an expanded cohort to identify additional variants and refine regional evidence. This connects broad discovery with focused validation, so researchers can prioritize candidate variants without repeating a full-genome survey.
Known genes, promoters, enhancers, or conserved elements can be studied together even when they are dispersed across the genome. A single coordinated design keeps the target definition consistent across samples, which is particularly useful for multi-population comparisons.
Concentrating reads on selected loci can increase evidence for low-frequency variants within those regions. The benefit is narrower, more reviewable candidate lists, although detection capability still depends on sequencing allocation, DNA quality, local sequence context, and cohort size.
Variants nominated in one population can be tested in additional populations, breeding materials, or geographic groups. This reveals whether candidate alleles are shared, population-specific, or associated with distinct regional haplotypes.
From chip to SNP: Rapid development and evaluation of a targeted capture genotyping-by-sequencing approach to support research and management of a plaguing rodent
Journal: PLOS ONE
Published: 2023
Oh KP, Van de Weyer N, Ruscoe WA, Henry S, Brown PR. From chip to SNP: Rapid development and evaluation of a targeted capture genotyping-by-sequencing approach to support research and management of a plaguing rodent. PLOS ONE. 2023;18(8):e0288701.
The researchers needed a flexible SNP genotyping approach for wild house-mouse populations that could support population assignment, kinship analysis, and continued monitoring. An existing array contained many markers that were uninformative for the focal wild populations, creating a clear need for a population-specific target set.
Array data from mice sampled across southeastern Australia were used to select informative SNPs for a custom hybridization-capture panel. The final design targeted 3,651 SNPs, and the researchers applied it to 320 mice before conducting sequencing QC, genotype comparison, principal-component analysis, discriminant analysis, population assignment, and kinship inference.
After quality filtering, 317 samples and 3,565 targeted SNPs were retained. For 47 samples measured by both methods, 98.9% of called genotypes were concordant. The targeted dataset distinguished mice from two adjoining farms and achieved a mean population-assignment rate of 0.955 in the study's cross-validation procedure.
Figure 5: Original case-study summary of the cohort, custom target capture workflow, and population-genetic findings reported by Oh et al. (2023).
This study shows how a target set derived from prior discovery data can support focused population analysis in a larger follow-up cohort. It also illustrates an important boundary: a panel optimized for one population may require redesign before use in a different population, because marker informativeness and ascertainment bias can change.
Method selection should begin with the research question, not with a preferred technology. Targeted resequencing is usually the best fit when the regions are already known and sequence-level variant discovery must scale across a large cohort; broader or fixed-marker methods are more appropriate when the question falls outside that boundary.
| Decision Factor | Targeted Resequencing | WGS | WES | SNP Genotyping |
| Best starting point | Known genes, loci, intervals, or coordinates | Open-ended genome-wide discovery | Coding-exon discovery | Known, predefined markers |
| Variant scope | Sequence variants within custom targets | Genome-wide sequence variants | Variants mainly in annotated exons | Alleles represented on the marker set |
| Noncoding targets | Can be included by design, helping studies test regulatory hypotheses | Included genome-wide, supporting discovery beyond known regions | Usually limited, so it may miss a regulatory-region question | Only included when a marker was predefined |
| Data volume | Focused, reducing storage and review burden per sample | Largest, providing breadth at higher analysis burden | Intermediate, focused on coding sequence | Compact genotype matrix, but no de novo sequence discovery |
| Large-cohort role | Follow-up, validation, regional discovery, and population comparison | Broad discovery when targets are not yet known | Exonic discovery when coding variation is the priority | Efficient testing of an established marker set |
| Key limitation | Cannot interrogate regions excluded from the design | More data than a targeted question may require | Does not represent most noncoding regions | Cannot discover variants absent from the marker panel |
Best for: studies with a defensible gene or region list, a suitable reference assembly, and a need to examine the same targets across many individuals. The method is especially useful after WGS, GWAS, QTL mapping, or pilot sequencing has already identified where to look.
Not the best choice for: open-ended discovery when the relevant genomic regions are still unknown. In that situation, Whole Genome Re-sequencing for Population Genetics may be more appropriate. If the question is limited to coding exons, Whole Exome Sequencing for Population Genetics may provide a more standardized scope. If only established markers are required and new variant discovery is unnecessary, the SNP Genotyping Service may be a better fit.
Confidence in a population-scale targeted sequencing project comes from decisions that remain traceable from target selection through cohort-level delivery. CD Genomics connects target-design review, method selection, cohort QC, and bioinformatics scope, helping project teams understand not only what will be done, but why each decision matters for the final comparison.
Research-question-led target review: gene lists, GWAS loci, QTL intervals, and custom coordinates are reviewed against the reference assembly and intended analysis. This reduces the risk of building a technically valid design that does not answer the study's biological question.
Any platform, depth, input, or performance specification not confirmed for the individual project remains subject to feasibility review. This transparency is part of the service design: difficult targets and method limitations should be identified before they become unexplained gaps in a large cohort.
You can begin with gene symbols, transcript identifiers, genomic coordinates, a BED file, GWAS loci, QTL intervals, or prior variant results. Always include the reference assembly version and explain how each target relates to the study objective. This lets feasibility review detect coordinate mismatches, repetitive regions, or missing interval context before design.
There is no useful universal limit because feasibility depends on the number and length of intervals, genome complexity, repeats, GC composition, homology, sample count, and the targeting method. A compact amplicon design and a broad hybrid-capture design solve different problems, so target size should be evaluated together with the cohort and expected analysis.
Coverage is planned from the variant classes, expected allele frequencies, ploidy, target architecture, sample quality, and acceptable missingness. The key decision is not a generic depth number; it is whether the planned evidence can support the intended variant and cohort comparison. The project plan should therefore define both sequencing allocation and QC interpretation.
A suitable reference assembly is normally needed for coordinate-based design, read alignment, and variant annotation. If the focal species lacks a high-quality reference, a related assembly, transcriptome, prior amplicons, or de novo sequence resource may sometimes support a modified strategy, but uncertainty and cross-species divergence must be evaluated before design.
Yes. Existing WGS, WES, GWAS, QTL, VCF, or pilot results can be used to nominate targets when the reference assembly, coordinates, and supporting evidence are provided. External results are reviewed rather than assumed to be directly design-ready, which reduces errors caused by outdated assemblies or ambiguous feature definitions.
It can support allele-frequency summaries when sampling groups, genotype QC, and callable sites are defined consistently. Regional haplotype analysis may also be possible when marker density and phase information are adequate. These analyses should be planned before sequencing because cohort balance, metadata, and target spacing affect what can be inferred.
Use Variant Calling Service when the primary need is consistent variant processing from existing sequencing data. For discovery-stage population studies, explore the Genome-wide Association Analysis Service or QTL Location Analysis Service; their results can provide the loci and intervals for a later targeted resequencing study.
Next step: assemble the target list, reference assembly, sample count, population groups, existing discovery evidence, and desired outputs. These inputs allow a technical review to determine whether capture, amplicon targeting, WGS, WES, or SNP genotyping best matches the project.
References