How Can Public scRNA-Seq Data Reveal Immunotherapy Resistance Mechanisms? A Colorectal Cancer Case Study

Public single-cell RNA sequencing (scRNA-seq) datasets can support new biological discovery when the study question, cohort selection, analytical comparisons, and validation strategy are rigorously matched. Public archives now contain large numbers of single-cell datasets, but deposition does not mean every biological question has already been answered. Re-analysis can still add value when investigators introduce a focused mechanistic question, compare independent cohorts carefully, or examine cell states and regulatory programs that were outside the scope of the original publication. Cross-study synthesis has also shown that recurrent transcriptional programs can emerge only after many independent tumor datasets are evaluated together [6].
Generating credible insights from public data requires moving beyond descriptive re-clustering. Recreating UMAP plots or listing generic differential markers is not, by itself, a mechanistic study. This guide uses a recent colorectal cancer case study to show how public scRNA-seq data can be moved through a question-driven sequence—from broad cell-landscape screening to T-cell state analysis, regulatory and communication hypotheses, and independent tissue validation. The central study examined anti-PD-1 resistance in mismatch repair-deficient (dMMR) / microsatellite instability-high (MSI-H) colorectal cancer [1]. This guide discusses research-use discovery workflows, computational study design, and preclinical validation principles rather than clinical diagnosis, patient management, or treatment selection.
To identify appropriate public cohorts for your own research questions, explore our curated selection guide on 26 Single-Cell and Spatial Omics Atlases Worth Bookmarking. For validation strategies, consult our framework on How To Validate Single-Cell RNA-Seq Data?.
Key Takeaways:
- Question-Driven Re-Analysis: The value of public single-cell data depends on formulating a precise, testable biological question rather than maximizing dataset size.
- Hierarchical Analytical Funnel: Successful data mining begins with an unbiased, global microenvironmental survey to nominate candidate cell compartments before narrowing to specialized subsets.
- Complementary Analytical Modules: Sub-clustering, GSEA pathway enrichment, SCENIC transcription-factor regulons, and CellChat communication analysis address distinct biological layers and cannot substitute for one another.
- Hypothesis vs Proof: Computational gene regulatory networks and ligand-receptor interactions generate prioritized mechanistic hypotheses that require orthogonal biological validation.
- Independent Multi-Tier Validation: Transitioning computational candidates into credible research findings requires independent tissue cohorts, protein-level assays (e.g., IHC), and survival association modeling.
Can Public scRNA-Seq Data Still Support New Biological Discovery?
A prevalent misconception in functional genomics is that once a single-cell dataset has been deposited in an archive, its informational value has been completely extracted. In practice, primary single-cell studies typically focus on broad cell-atlas generation or address a specific phenotypic contrast predefined by the original authors. Vast layers of granular biological information—such as rare transitional states, alternative regulon networks, cell-cell communication rewiring, and subclonal immune dynamics—frequently remain unexamined.
Public single-cell data re-analysis yields three distinct forms of new scientific value:
- Novel Mechanistic Interrogation: Applying a focused biological hypothesis (e.g., specific therapeutic resistance programs, metabolic checkpoints) to existing high-quality datasets that lacked such framing in the primary study.
- Cross-Cohort Harmonization & Meta-Analysis: Integrating independent public cohorts can test whether candidate programs recur across studies and reduce reliance on cohort-specific observations, provided biological and technical differences are explicitly modeled [6].
- In Silico Discovery Paired with Targeted Validation: Public data can serve as a hypothesis-generating discovery engine that prioritizes specific candidate genes, pathways, or cell states before focused orthogonal validation in independent biological material.
The decisive factor distinguishing impactful data mining from superficial re-analysis is methodological discipline. Meaningful discovery requires an evidence-based progression: establishing a clear biological question, executing rigorous cohort curation, applying multi-layered bioinformatics, and closing the discovery loop through independent experimental validation.
The Biological Question: Why Do Some MSI-H Colorectal Cancers Resist Anti-PD-1 Therapy?
The value of disciplined data mining is illustrated by an oncology case study focused on immunotherapy resistance in dMMR / MSI-H colorectal cancer. These tumors can respond to immune checkpoint blockade targeting the PD-1/PD-L1 axis, yet a subset of patients remains resistant and the underlying cellular mechanisms are heterogeneous. The study therefore asked whether cell-resolved transcriptional states and T-cell interaction programs could help explain differences between anti-PD-1-sensitive and -resistant tumors [1].
Rather than generating a new single-cell discovery cohort, the investigators re-analyzed a public scRNA-seq dataset from NCBI SRA accession PRJNA932556, comprising six MSI-H colorectal cancer samples categorized as three anti-PD-1-sensitive and three resistant tumors. After quality control, the study analyzed 43,335 cells and used the public cohort to define the focused research question: Which T-cell states, regulatory programs, and intercellular communication patterns are associated with resistance to anti-PD-1 therapy in dMMR/MSI-H colorectal cancer? [1].
Stage 1: Profile the Complete Cellular Landscape Before Selecting a Target Mechanism
A common error in single-cell data mining is premature narrowing—jumping directly to a favorite gene or cell type based on literature bias. A rigorous workflow begins with an unbiased, global survey of the complete tumor microenvironment (TME).
In the case study, researchers converted the public count matrix into a Seurat object and applied study-specific quality filters, including molecule-count, detected-gene, mitochondrial-signal, and low-frequency gene criteria before downstream integration. The exact thresholds belonged to this dataset and should not be treated as universal scRNA-seq QC rules. The investigators then normalized the data, selected variable genes, identified integration anchors, and compared cell populations across response groups [1].
Unsupervised analysis resolved 11 major cell types across 43,335 cells. The overall cellular composition of sensitive and resistant tumors was broadly comparable, although monocytes increased and intestinal epithelial and clustered-cell populations decreased in the resistant group. Endothelial cells and monocytes showed the largest numbers of differentially expressed genes at the broad cell-type level. T cells were the predominant population, but their aggregate expression profile was not the most globally altered; instead, several T-cell subclusters showed pronounced transcriptional changes. This within-lineage heterogeneity motivated the focused T-cell analysis that followed [1].
Figure 2. Global microenvironmental screening: evaluating lineage distributions across treatment-response groups to isolate candidate compartments for deep subpopulation analysis.
Stage 2: Resolve T-Cell Functional Subsets and Pathway Programs
Once the T-cell compartment was identified as the key divergent population, investigators isolated all T cells and performed iterative high-resolution sub-clustering. Coarse lineage definitions (CD4+ vs CD8+) were insufficient to capture resistance dynamics; fine-grained functional annotation was required.
Sub-clustering identified several specialized T-cell states:
- Exhausted T Cells (Tex): Characterized by sustained expression of inhibitory immune checkpoints (e.g., PDCD1, HAVCR2, LAG3).
- Cytotoxic Effector T Cells (GZMK+ T): Defined by granzyme K and intermediate effector differentiation.
- Stress-Response T Cells (TSTR): Enriched for heat-shock and proteotoxic stress response genes (e.g., HSP90AA1, HSPA1A).
- Regulatory T Cells (Treg): Marked by canonical FOXP3 and immunosuppressive cytokine expression.
- Gamma-Delta T Cells (γδ T): Expressing alternative TCR chains with innate-like cytotoxic properties.
Gene Set Enrichment Analysis (GSEA) showed that resistance-associated programs differed by T-cell subset rather than following a single shared pathway. In exhausted T cells (Tex), antigen-processing and presentation programs were enriched in resistant tumors, while oxidative phosphorylation-related programs were reduced. GZMK+ T cells also showed resistance-associated changes involving antigen processing, natural killer cell differentiation, interferon-related programs, and stress-response pathways. In TSTR cells, protein-folding and molecular-chaperone programs were inhibited in the resistant group, whereas gamma-delta T cells in the sensitive group showed activation of protein-folding programs. These results support a model in which treatment outcome is associated with distinct functional states across multiple T-cell subsets rather than simply the presence or absence of T cells [1].
Stage 3: Infer Regulatory Networks and Intercellular Communication Hypotheses
To move from describing cell states to identifying candidate upstream drivers, the investigators deployed two advanced computational frameworks: Gene Regulatory Network (GRN) analysis using SCENIC, and cell-cell communication mapping using CellChat.
1. Single-Cell Regulatory Network Inference (SCENIC)
SCENIC combines single-cell co-expression with transcription-factor motif information to infer regulons and compare regulatory activity across cell states [2]. In the CRC case study, SCENIC identified substantial regulatory-network dysregulation in TSTR and gamma-delta T cells. Differential transcription-factor programs included KLF6, ATF3, FOS, JUN, and NR3C1, nominating these factors and their target networks for further mechanistic investigation. The study did not establish that a single AP-1 program was uniformly activated across resistant tumors; rather, it reported subset-specific transcription-factor dysregulation that requires functional follow-up [1].
2. Intercellular Communication Analysis (CellChat)
CellChat infers and compares candidate ligand-receptor communication networks from single-cell transcriptomic data; it does not directly measure physical signaling events [3]. In the case study, resistant tumors showed increased numbers and strengths of inferred interactions among T-cell subpopulations, together with altered signaling patterns. The CD69-KLRB1 axis was notably reduced in resistant specimens, making it a candidate communication feature associated with treatment response rather than a proven causal mechanism [1].
| Analytical Module | Specific Research Question Addressed | Biological Evidence Generated | Key Boundary: What It Cannot Prove |
|---|---|---|---|
| Sub-Clustering & Re-Annotation | Which specific cell sub-states differ between response groups? | Resolves specialized subsets (Tex, TSTR, γδ T) hidden in coarse clusters | Does not establish whether cell state shifts are a cause or consequence of resistance |
| GSEA Pathway Enrichment | What biological programs and metabolic pathways are altered? | Statistical enrichment of oxidative stress vs protein folding programs | Pathways reflect correlated gene sets; does not confirm physical enzyme activity |
| SCENIC Regulon Inference | Which master transcription factors coordinate the resistance signature? | Identifies candidate TF drivers (e.g., FOS/JUN AP-1 regulons) with motif support | Inferred from transcript correlation; does not verify physical TF-DNA binding in vivo |
| CellChat Communication | How does cell-cell signaling rewiring alter microenvironmental crosstalk? | Pinpoints candidate ligand-receptor downregulation (e.g., CD69–KLRB1) | Computational inference from mRNA; does not prove protein-level physical contact |
Figure 3. Analytical progression: linking single-cell sub-clustering with differential pathway enrichment, master regulon modeling, and ligand-receptor communication mapping.
Stage 4: Prioritize Computational Candidates for Independent Tissue Validation
The definitive strength of the case study lies in avoiding a purely computational conclusion. The authors utilized their bioinformatic findings as a screening funnel to prioritize specific molecular candidates—FOS and KLRB1—and validated them in an independent clinical research cohort.
The investigators evaluated prospectively collected tumor and paired adjacent normal tissue specimens from 190 patients with colorectal cancer. Immunohistochemistry was used to assess FOS and KLRB1 protein expression, and the study related those measurements to clinicopathological variables and survival outcomes [1].
- FOS Protein Validation: In this study cohort, higher FOS protein expression in tumor tissue was associated with more aggressive clinicopathological features and poorer survival outcomes [1].
- KLRB1 Protein Validation: KLRB1 protein expression was lower in tumor tissue than in paired adjacent normal tissue. Within the analyzed cohort, higher KLRB1 expression was associated with longer overall survival (OS, P = 0.024) and progression-free survival (PFS, P = 0.027) [1].
- Multivariate Risk Modeling: Cox regression in the study supported an independent prognostic association for KLRB1 expression. These results are cohort-specific and should be interpreted as research evidence requiring additional independent and functional validation rather than as a validated clinical test [1].
By connecting public scRNA-seq screening to an independent 190-patient tissue cohort, the investigators moved from computational prioritization to orthogonal protein-level and outcome-associated validation. The article was published online on September 9, 2025 and appears in the 2026 volume of the International Journal of Surgery [1].
Figure 4. Translational validation workflow: prioritizing in silico candidate genes for multi-modal validation in independent patient tissue cohorts.
A Reusable Blueprint for Public Single-Cell Data Mining
Researchers can replicate this study architecture across diverse disease models, developmental systems, and pharmacology studies by following a structured, seven-step blueprint:
- Step 1: Define a Precise Biological Question: Focus on a specific phenotypic contrast, therapeutic resistance phenotype, or cellular transition with clear translational relevance.
- Step 2: Curate and Screen Public Datasets: Filter public repositories for cohorts that provide high-quality raw data paired with complete, transparent clinical or experimental metadata.
- Step 3: Harmonize Data and Account for Batch Effects: Execute unified quality filtering, remove computational multiplets, and harmonize data across donors and experimental batches.
- Step 4: Execute Broad Landscape Screening: Characterize global cell lineage distributions across experimental groups to objectively nominate the most relevant compartment.
- Step 5: Perform Focused Deep Sub-Clustering: Isolate target compartments for high-resolution sub-clustering, functional marker annotation, and GSEA pathway enrichment.
- Step 6: Model Regulatory Drivers and Intercellular Crosstalk: Deploy SCENIC to nominate master transcription factors and CellChat to map cell-cell communication rewiring.
- Step 7: Formulate an Independent Validation Strategy: Prioritize top candidate markers for validation in independent tissue cohorts using orthogonal modalities (e.g., multiplex IHC, spatial transcriptomics, targeted qPCR).
To design and execute customized single-cell mining workflows, explore our dedicated Single-Cell Omics Bioinformatics and Data Mining Service and our specialized Intercellular Communication Analysis Services.
Common Pitfalls in Public scRNA-Seq Data Mining
Re-analyzing public datasets carries distinct technical and methodological risks. Investigators should actively guard against common failure modes:
- Treating Single Cells as Independent Biological Replicates: Treating cells from the same donor as independent statistical replicates inflates significance and can produce false discoveries. Prefer donor-aware approaches such as pseudobulk aggregation or models that account for biological replicate structure [4].
- Confounding Treatment Groups with Experimental Batches: If all "sensitive" samples were processed in Study A and all "resistant" samples in Study B, observed differences will reflect technical batch variation rather than biology. Verify that comparisons are balanced across batches.
- Over-Interpreting Computational Inferences: Treating SCENIC regulons or CellChat ligand-receptor predictions as definitive proof of biochemical mechanism is a critical error. Computational predictions provide prioritized hypotheses, not functional proof.
- Selective Hypothesis Confirmation: Repeatedly changing clustering parameters, subgroup definitions, or pathway filters until they support a preferred story increases confirmation bias. Predefine core comparisons when possible, document analytical decisions, and test whether key findings are robust to reasonable alternative settings.
- Skipping Independent Orthogonal Validation: Reporting a "novel biomarker" based solely on public computational re-analysis without validating expression in an independent biological cohort substantially undermines study reproducibility.
When Public Data Are Not Enough
While public single-cell re-analysis is highly cost-effective, investigators will encounter scenarios where public data cannot resolve the hypothesis:
- Critical Clinical Annotations Are Missing: Public repositories frequently omit crucial metadata—such as exact drug regimens, timing of biopsy collection, or long-term clinical endpoints—preventing rigorous group stratification.
- Target Cell Populations Are Depleted by Dissociation: Enzymatic tissue dissociation can selectively destroy fragile cell types (e.g., mature adipocytes, plasma cells, specific stromal lineages), leaving them underrepresented in public scRNA-seq suspensions.
- Spatial Tissue Context Is Mandatory: When a hypothesis depends on tumor-stroma boundaries, immune hubs, tertiary lymphoid architecture, or invasive fronts, dissociated single-cell data cannot directly preserve tissue location. Spatial profiling has shown that colorectal cancer immune programs can be organized into localized multicellular niches [5]. In such cases, generating de novo spatial data may be needed; explore our strategic framework on Spatial Transcriptomics Cancer Research: 4 Core Strategies.
- Underpowered Donor Replicates: Public datasets often contain only a handful of patient donors per group, precluding robust statistical generalization across heterogeneous disease populations.
When public datasets prove insufficient, generating de novo single-cell or spatial sequencing under controlled experimental protocols becomes essential. Learn more about our experimental platforms at Single-Cell Sequencing Services and routine processing through Single-Cell RNA-Seq Data Analysis Service.
FAQs
From Public Data to a Project-Specific Validation Plan With CD Genomics
Public-data mining is most useful when the research question determines the dataset selection, comparison design, analytical depth, and validation path. CD Genomics supports research-use public dataset curation, cross-study integration, cell-state and regulatory analysis, communication analysis, and validation prioritization through our Single-Cell Omics Bioinformatics and Data Mining Service. When public cohorts cannot represent the required condition, cell population, or tissue context, project-specific single-cell or spatial profiling can be considered as a complementary research strategy.
Research Use and Trust Statement
This article summarizes published research and methodological considerations for research-use and preclinical study design. Study-specific associations should not be interpreted as validated diagnostic, prognostic, or treatment-selection claims. Computationally inferred regulatory or communication relationships require appropriate independent or functional validation.
References
- Chen Y, Liu T, Min G, et al. Single-cell transcriptome analysis reveals regulatory programs associated with tumor resistance during immunotherapy in colorectal cancer. International Journal of Surgery. 2026;112(1):694–708. Published online September 9, 2025.
- Aibar S, González-Blas CB, Moerman T, et al. SCENIC: single-cell regulatory network inference and clustering. Nature Methods. 2017;14(11):1083–1086.
- Jin S, Plikus MV, Nie Q. CellChat for systematic analysis of cell–cell communication from single-cell transcriptomics. Nature Protocols. 2025;20:180–219. Published online September 16, 2024.
- Squair JW, Gautier M, Kathe C, et al. Confronting false discoveries in single-cell differential expression. Nature Communications. 2021;12:5692.
- Pelka K, Hofree M, Chen JH, et al. Spatially organized multicellular immune hubs in human colorectal cancer. Cell. 2021;184(18):4734–4752.e20.
- Gavish A, Tyler M, Greenwald AC, et al. Hallmarks of transcriptional intratumour heterogeneity across a thousand tumours. Nature. 2023;618(7965):598–606.