Aggregate rare variants into defensible gene and region tests with documented masks, burden and kernel methods, cohort-aware modeling, and traceable candidates.
Rare variants are observed too infrequently for many single-variant association tests to achieve useful power. Rare variant association analysis addresses this problem by grouping variants within a gene, transcript, functional element, pathway, or predefined region and testing their combined relationship with a phenotype. The result is a set-based signal that can reveal aggregated genetic effects while reducing the number of tests relative to evaluating every rare variant separately.
CD Genomics supports WES and WGS cohorts, as well as external VCF, phenotype, and covariate data, with cohort QC, variant annotation, frequency and functional filtering, grouping, burden tests, SKAT, SKAT-O, covariate adjustment, population-structure or relatedness handling, multiple-testing control, visualization, and candidate prioritization. Analysis choices are documented because no single test is uniformly most powerful across all genetic architectures.
Figure 1: Rare variant association analysis converts sparse variant observations into predefined set-level tests.
A rare allele may occur in only a few participants, leaving too little information for a stable single-site estimate. Aggregation increases the number of informative carriers by testing multiple qualifying variants together. The gain depends on the grouping rule, allele-frequency threshold, functional annotation, causal-variant proportion, effect directions, sample size, phenotype distribution, and ancestry composition.
Variant sets should be defined before association results are inspected. Common units include genes, exons, protein domains, enhancers, sliding windows, and pathways. Multiple masks may capture different biological assumptions—for example, protein-truncating variants alone versus protein-truncating plus predicted damaging missense variants—but every additional mask increases the testing burden and should have a stated rationale.
| Method | Working Assumption | When It Can Be Useful |
| Burden test | Eligible variants influence the phenotype in a broadly consistent direction; collapsing or weighting creates one burden variable per set. | Sets enriched for likely functional variants with a relatively high causal proportion and aligned effects. |
| SKAT | Variant effects may differ in magnitude or direction, and some eligible variants may be neutral. | Heterogeneous sets where protective and risk effects may coexist or only a subset of variants is causal. |
| SKAT-O | The optimal balance between burden-like and variance-component evidence is not known beforehand. | Broad scans where genetic architecture varies among genes and an adaptive omnibus test is appropriate. |
The test is only one component of the analysis. Variant selection, weights, null model, phenotype type, covariates, relatedness, case-control imbalance, and sample size can change calibration and power. Sensitivity analyses across justified masks or frequency thresholds can show whether a signal is robust or driven by one modeling choice.
Figure 2: Test selection follows the expected proportion and direction of causal effects, while SKAT-O adapts across a range of architectures.
When sequencing is required, related entry points include Whole Exome Sequencing and Whole Genome Resequencing. Existing sequence data can first undergo Variant Calling and cohort harmonization.
| Project Situation | Recommended Path | Key Consideration |
| Jointly processed WES/WGS variants with phenotype and covariate data | Direct entry: annotate, filter, construct sets, and test | Genotype quality, ancestry, and callability are verified before testing. |
| Existing cohort VCF with documented provenance | Variant QC and harmonization before association testing | Cross-processed cases and controls may need recalling or harmonization. |
| Raw FASTQ or BAM files only | Variant Calling followed by this analysis | Consistent calling across the cohort is required for set-based tests. |
| Sequencing not yet performed | Whole Exome Sequencing or Whole Genome Resequencing then analysis | Depth, cohort size, and study design should be aligned with the planned tests. |
| Small cohort or very few expected carriers | Power and expected-carrier evaluation before analysis | Sparse data may require SKAT-O or simulation-based feasibility review. |
| Phenotype, covariates, or relatedness information incomplete | Complete metadata before starting | Underspecified phenotypes and missing covariates degrade every downstream test. |
1. Analysis specification
We define the phenotype, cohort, hypothesis, grouping units, MAF thresholds, functional masks, covariates, relatedness strategy, tests, and multiple-testing plan before association analysis.
2. Sample and genotype QC
Sample identity, sex checks when appropriate, duplicates, missingness, depth or genotype quality, ancestry, relatedness, and batch are reviewed. Variant QC is aligned across groups.
3. Annotation and rare-variant filtering
Variants are annotated against the agreed assembly and transcript set, then filtered using cohort and external frequencies, consequence classes, callability, and project-specific evidence.
4. Set and mask construction
Eligible variants are grouped into genes or other predefined regions. Set sizes, carrier counts, and variants contributing to each mask are recorded.
5. Null model and association testing
Phenotype-appropriate models incorporate covariates, ancestry, and relatedness as required. Burden, SKAT, SKAT-O, or agreed methods are applied with calibrated handling of cohort structure.
6. Multiple testing and sensitivity review
Family-wise or false-discovery procedures are applied to the defined test family. Signals are compared across masks, methods, ancestry groups, or leave-one-variant-out checks when justified.
7. Prioritization and delivery
Significant and suggestive sets are linked to contributing variants, annotations, effect summaries, QC, plots, and a methods report for downstream research.
Figure 3: Rare variant analysis is specified, tested, and reviewed as one connected set-based workflow rather than a single test.
| Parameter | Typical Scope | Review Point | Why It Matters |
| Genotype input | Jointly processed WES/WGS variants or a documented cohort VCF | Genotype-level quality fields, reference assembly, and callability | Comparable variants across the cohort are the foundation of set tests. |
| Phenotype and covariates | Coded binary or quantitative traits; age, sex, site, batch, ancestry components | Phenotype definition, missingness, and case-control balance | Underspecified phenotypes and omitted covariates distort calibration. |
| Variant filtering | MAF thresholds, consequence classes, and functional scores | Cohort and external frequencies, callability, and evidence sources | Prespecified thresholds keep filtering rules transparent. |
| Set and mask construction | Gene, region, or pathway units with one or more justified masks | Set sizes, carrier counts, and stated rationale per mask | Every additional mask increases testing burden and needs a reason. |
| Testing strategy | Burden, SKAT, SKAT-O, or agreed methods | Null model, covariates, relatedness, and ancestry handling | No single test is uniformly most powerful across genetic architectures. |
| Multiple testing and sensitivity | Family-wise or false-discovery procedures; mask and method comparisons | Leave-one-variant-out, ancestry strata, and convergence checks | Sensitivity review shows whether a signal depends on one modeling choice. |
| Deliverables | QC summaries, annotated variant table, set definitions, result tables, figures, prioritized candidates, report | Effect, p-value, multiple-testing, carrier-count, and convergence fields | Traceable candidates let downstream research verify every association. |
Rare variants are often geographically localized, making ancestry confounding especially important. Principal components or other population variables may be included, cohorts may be stratified and meta-analyzed, or mixed-model approaches may be selected. Related participants cannot simply be treated as independent observations because this may inflate test statistics. Extreme case-control imbalance and very small carrier counts also require methods and interpretation suited to sparse data.
Population Structure Analysis can support ancestry review, while Genome-wide Association Analysis provides complementary common-variant evidence. Rare variant analysis remains a distinct set-based workflow rather than an extra filter applied to a GWAS result.
Figure 4: Representative outputs connect gene-level signals to model calibration, masks, carrier counts, and contributing variants.
| Best For | Not For |
| WES/WGS case-control or quantitative-trait cohorts, gene-based discovery, pharmaceutical target and preclinical biomarker research, and external VCF plus phenotype projects with sufficient cohort information | Underspecified phenotypes, unharmonized case/control sequencing, variant sets defined after inspecting association results, or cohorts too small to support the planned set tests |
The impact of rare germline variants on human somatic mutation processes
Journal: Nature Communications
Published: 2022
Vali-Pour M, Park S, Espinosa-Carrasco J, Ortiz-Martínez D, Lehner B, Supek F. The impact of rare germline variants on human somatic mutation processes. Nature Communications. 2022;13:3724.
The authors investigated whether inherited rare variants influence the rates and patterns of somatic mutation observed among cancer genomes.
They performed gene-based rare variant association analyses across human cancer sequencing cohorts, using prioritized rare putative loss-of-function variant sets, multiple inheritance models, and a combined burden and variance test. Discovery results were retested in an independent validation cohort.
The study reported 207 associations that replicated at an empirical false discovery rate of 1%, spanning 42 genes and 15 somatic mutational phenotypes. The results show how predefined masks, combined tests, multiple-testing control, and independent validation can be linked in a rare-variant workflow.
Figure 5: Original case-study summary of rare variant grouping, association testing, and independent validation reported by Vali-Pour et al. (2022).
The study illustrates a rigorous pattern for rare variant research: biologically defined variant sets, models that accommodate different genetic architectures, explicit multiple testing, and independent replication. Its findings are specific to the analyzed cohorts and phenotypes.
CD Genomics connects cohort QC, annotation, variant-set construction, statistical modeling, and candidate review so that a gene-level p-value can be traced back to the samples, variants, masks, and assumptions that produced it.
Predefined analysis specifications: frequency thresholds, annotation masks, grouping rules, tests, covariates, and correction procedures are documented before result review.
There is no universal threshold. The cutoff should reflect sample size, ancestry, external frequency resources, expected genetic architecture, and method. One or more prespecified thresholds may be analyzed as separate masks.
Burden tests favor aligned effects and a high causal proportion; SKAT tolerates mixed directions and neutral variants; SKAT-O adapts between them. Choice depends on biology, mask specificity, cohort size, and phenotype.
Yes, if genotype fields, reference assembly, calling workflow, sample metadata, phenotype, covariates, and QC provenance are sufficient. Recalling or harmonization may be recommended when cases and controls were processed differently.
Options include ancestry covariates, stratified analysis and meta-analysis, mixed models, and relationship matrices. The appropriate approach depends on cohort structure and the selected association method.
Power depends on phenotype prevalence or variance, carrier counts, set size, causal proportion, effect magnitude and direction, ancestry, covariates, and testing burden. Feasibility should be evaluated through expected carriers and power simulations when possible.
Generate or process cohort data through Whole Exome Sequencing, Whole Genome Resequencing, and Variant Calling. Complement gene-based tests with Genome-wide Association Analysis and Population Structure Analysis.
References