Single-Cell Omics Bioinformatics and Data Mining Service

Single-cell and spatial datasets often contain more biological evidence than a routine processing report can reveal. Our single-cell omics bioinformatics and data mining service turns public or customer-generated datasets into a question-driven evidence framework for cell states, regulatory mechanisms, interactions, biomarkers, and validation priorities.

Use the service to:

  • Find and curate datasets that directly support your research question
  • Harmonize studies, platforms, cohorts, and metadata for defensible comparison
  • Move beyond routine clustering into mechanism-focused analysis
  • Generate reproducible figures, tables, code, and validation hypotheses

Discuss a Data Mining Project

Single-cell omics bioinformatics and data mining

Question-Driven Analysis Beyond a Routine Pipeline

Routine analysis commonly produces QC, clustering, marker genes, and basic cell annotations. Data mining begins with a different question: what evidence is required to support or reject a biological model? We translate that question into an evidence map, identify suitable datasets, define comparison boundaries, and select analysis modules only when they contribute to the primary claim.

Evidence Mapping

Connect each biological claim to required datasets, contrasts, analyses, and validation.

Data Curation

Assess provenance, metadata, quality, comparability, and potential confounders.

Deep Analysis

Interrogate cell states, trajectories, interactions, regulation, and cross-study patterns.

Research Delivery

Provide interpretable results, reproducible methods, and practical validation priorities.

Supported single-cell and spatial omics datasets Figure 2. Public and customer datasets can be combined when their designs support the same research question.

For standard processing of a new scRNA-seq dataset, see our Single-Cell RNA-Seq Data Analysis Service. This service is intended for deeper, integrative, or multi-dataset questions.

Supported Single-Cell, Spatial, and Cohort Data

Public Data Sources

Projects can incorporate suitable studies from repositories such as GEO, ENA, ArrayExpress, NGDC, the Human Cell Atlas, Zenodo, and spatial-data resources. Dataset inclusion is based on biological relevance, raw or processed data availability, metadata completeness, and technical compatibility.

Customer-Generated Inputs

  • FASTQ or BAM files and count matrices
  • Seurat, AnnData, loom, or Cell Ranger outputs
  • Single-cell RNA, ATAC, immune repertoire, and multimodal data
  • Spatial matrices, coordinates, images, and segmentation files
  • Matched bulk, clinical, experimental, or cohort metadata

A data inventory is completed before analysis. Missing sample annotations, inconsistent identifiers, or unbalanced study designs are documented as evidence limitations rather than hidden during integration.

The quantities below are recommended planning targets rather than universal acceptance thresholds. A project can still be feasible with fewer samples when the study is exploratory, rare, or explicitly designed as a case series.

Data PackageMinimum / Recommended InputRequired Information
Raw sequencing dataComplete FASTQ or BAM files for 100% of included samples; all lanes and read groups suppliedSample-to-file map, library type, sequencing platform, reference genome, chemistry, and batch
Processed matricesAt least 1 complete gene-by-cell, peak-by-cell, clonotype, or spatial matrix per datasetFeature IDs, barcode IDs, filtering history, normalization status, and cell metadata
Analysis objects1 Seurat, AnnData, loom, or Cell Ranger output per dataset, with the matching software versionRaw-count layer retained where possible; embeddings alone are insufficient for reanalysis
Group comparisonAt least 3 biological samples per group are recommended when donor-level inference is requiredDonor, condition, tissue, time point, treatment, sex, age, and batch fields as applicable
Cross-study validationAt least 2 independent datasets or cohorts are preferred; reserve 1 independent validation cohort when availableInclusion and exclusion criteria must be defined before model or signature evaluation
Spatial integrationExpression matrix + x–y coordinates + tissue image + segmentation or spot geometry for every sectionPixel-to-coordinate scale, orientation, section identity, and region annotations

Sequencing Platforms and Analysis Environment

ItemSupported Scope
Single-cell platforms10x Genomics Chromium, BD Rhapsody, SMART-seq, SeekOne, and other documented scRNA-seq, scATAC-seq, CITE-seq, or TCR/BCR workflows
Spatial platformsStereo-seq, 10x Visium / Visium HD, Xenium, CosMx, and other formats with complete coordinate and image metadata
New sequencingNot required for a public-data or existing-data mining project; any new sequencing need is scoped as a separate experimental service
Primary softwareR and Python workflows using project-appropriate releases of Seurat, Scanpy, Bioconductor, Cell Ranger outputs, and modality-specific tools
ReproducibilitySoftware versions, parameters, random seeds, environment files, and analysis scripts documented in the delivery package

Review Your Data Package

Single-Cell Omics Data Mining Workflow

  1. Research Question and Evidence Map

    Define the biological claim, comparison, expected evidence, confounders, and validation route.

  2. Dataset Discovery or Upload

    Search public resources or inventory customer files, formats, metadata, and study provenance.

  3. QC, Curation, and Harmonization

    Review quality, annotate covariates, resolve identifiers, and select an appropriate integration strategy.

  4. Deep Analysis and Integration

    Apply approved cell-state, trajectory, interaction, regulatory, repertoire, spatial, or multi-omics modules.

  5. Biological Interpretation and Validation Plan

    Evaluate alternative explanations and prioritize findings for independent testing.

  6. Publication-Ready Delivery

    Deliver documented methods, reproducible code, figures, tables, and a structured research report.

Single-cell omics data mining workflow Figure 3. A traceable workflow from question definition to reproducible delivery.

Select Analysis Modules That Support the Claim

Analysis modules are selected from the biological claim and available evidence rather than applied as a fixed checklist. After curation and harmonization, the workflow branches into the analyses needed to test cell-state, transition, interaction, regulatory, repertoire, cross-study, or spatial hypotheses.

Cell State and Mechanism

  • Curation and annotation: quality review, harmonization, reference mapping, cell typing, and rare-population assessment.
  • Differential programs: pseudobulk or cell-level comparisons, pathways, signatures, and covariate-aware models.
  • State transitions: trajectory, pseudotime, lineage-associated programs, and dynamic gene modules.
  • Cell interactions: ligand–receptor inference, niche-specific signals, and cross-method consensus.

Integration and Validation

  • Gene regulation: transcription-factor activity, regulons, chromatin accessibility, and RNA–ATAC integration.
  • Immune repertoire: clonotypes, expansion, diversity, receptor–phenotype relationships, and group comparisons.
  • Cross-study meta-analysis: atlas mapping, conserved cell states, heterogeneity, and sensitivity analysis.
  • Spatial integration: cell-type mapping, neighborhood analysis, spatial domains, and tissue-context validation.

Projects combining chromatin accessibility and expression can use our Single-Cell ATAC + RNA-seq Service. Immune-focused studies can connect repertoire evidence through Single-Cell Immune Repertoire Sequencing.

Single-cell omics bioinformatics analysis workflow Figure 4. Question-dependent analysis branches connect curated datasets with cell-state, trajectory, interaction, regulatory, repertoire, cross-study, and spatial evidence.

Reproducible and Interpretation-Ready Deliverables

Representative single-cell data mining outputs Figure 5. Representative atlas, trajectory, interaction, and regulatory outputs.

  • Dataset inventory, provenance table, and inclusion rationale
  • QC, curation, and harmonized analysis objects
  • Processed matrices, metadata, annotations, and result tables
  • High-resolution figures with editable source data where applicable
  • Methods, software versions, parameters, and reproducible scripts
  • Biological interpretation, evidence limitations, and validation priorities
  • Project-specific analysis report and presentation-ready result summary

Research Questions This Service Can Help You Answer

Oncology and Tumor Microenvironment

Candidate tumor or stromal cell states found in one cohort may not recur across cancer types or datasets. Cross-study integration quantifies their prevalence, transcriptional programs, neighboring cells, and treatment or outcome associations, helping you distinguish reproducible tumor ecosystems from cohort-specific observations.

Immunology and Immune Repertoire

When a response appears to arise from only part of the immune compartment, joint phenotype and repertoire analysis connects clonal expansion with activation, exhaustion, tissue residency, and group-specific immune states. This helps you identify which clonotypes and functional states are most relevant to the response. See our Spatial Omics Solutions for Immunology.

Neuroscience

Disease-associated cells can vary across brain regions, donors, and published studies. Atlas mapping and differential-state analysis separate conserved populations from region-, donor-, or protocol-specific effects, helping you decide which cellular programs warrant experimental validation.

Development and Regeneration

To test whether a proposed progenitor population generates a mature lineage, trajectory, regulon, and spatial evidence can be evaluated together. The combined result prioritizes plausible transitions, regulatory drivers, and stage-specific markers for lineage-tracing or perturbation experiments.

Biomarker and Target Discovery

Biomarkers intended for heterogeneous patient populations need evidence beyond a single dataset. Meta-analysis tests cell-type specificity, prevalence, co-expression, and cross-cohort robustness, helping you prioritize candidates before committing resources to experimental validation.

Drug Response and Mechanism

Treatment may change overall expression while leaving the responding cell population unresolved. Cell-state and interaction analyses localize the response, identify resistant niches, and prioritize pathway or ligand–receptor hypotheses, supporting follow-up mechanism and combination-strategy studies.

Case Study: A Pan-Cancer Single-Cell Ecosystem Atlas

Source: Systematic dissection of tumor-normal single-cell ecosystems across a thousand tumors of 30 cancer types

Background

Single cancer cohorts often lack the scale needed to distinguish recurrent tumor-associated cell states from organ-, study-, or protocol-specific effects. The study assembled a pan-cancer evidence framework to identify reproducible cellular programs and test whether they remained relevant in spatial tissue and treatment-response cohorts.

Methods

Kang and colleagues curated 104 single-cell RNA-seq datasets comprising approximately 4.9 million cells from 1,070 tumors and 493 normal samples across 30 cancer types and 999 donors. The meta-atlas was evaluated with 137 spatial transcriptomics samples from 11 cancer types, 8,887 TCGA bulk transcriptomes, and 1,261 checkpoint inhibitor-treated tumors.

The team harmonized cell annotations across studies, used an AND-gating strategy to identify recurrent tumor–normal programs, applied non-negative matrix factorization to resolve cell states, constructed cell-state co-occurrence networks, and deconvolved spatial data to test where candidate populations occurred. Logistic-regression meta-analysis was then used to test associations with immunotherapy response.

Results

The analysis separated AKR1C1-positive and WNT5A-positive inflammatory fibroblast states with different organ preferences, interaction patterns, and spatial neighborhoods. It identified an interferon-enriched tumor community containing LAMP3-positive dendritic cells, CCL19-positive fibroblasts, T follicular helper cells, and other antigen-presentation states. A derived tertiary lymphoid structure signature distinguished TLS from non-TLS spatial regions and was associated with favorable immunotherapy response across the 1,261 treated tumors.

Conclusion

The study showed how multi-cohort integration can separate recurrent cell ecosystems from cohort-specific signals and connect candidate states with spatial organization and treatment response. A comparable project can deliver a harmonized atlas, donor-aware state frequencies, conserved and cohort-specific programs, interaction networks, spatial validation maps, and independently tested signatures to determine whether a proposed state is reproducible, where it occurs, and whether its clinical association persists across datasets.

Illustrative reconstruction of pan-cancer single-cell data mining and validation Figure 6. Illustrative reconstruction of how multi-cohort integration can progress from recurrent cell-state discovery to spatial and response validation; not reproduced from the source paper.

When Data Mining Is the Right Next Step

Data mining is most useful when suitable datasets already exist and the main challenge is to test a biological claim across cells, donors, cohorts, modalities, or tissue contexts rather than generate another routine dataset.

Best for

  • Public-data discovery or customer-data reanalysis
  • Cross-study validation and evidence synthesis
  • Mechanism, biomarker, interaction, or regulatory questions

Consider another approach when

If the required cell population, condition, modality, or metadata is absent from available datasets, new data generation may be more appropriate. For routine processing of a newly generated dataset, use our Single-Cell RNA-Seq Data Analysis Service.

Choose the Right Analysis Scope

Frequently Asked Questions

References

  1. Stuart T, Butler A, Hoffman P, et al. Comprehensive Integration of Single-Cell Data. Cell. 2019;177:1888–1902.e21.
  2. Hao Y, Hao S, Andersen-Nissen E, et al. Integrated analysis of multimodal single-cell data. Cell. 2021;184:3573–3587.e29.
  3. Luecken MD, Büttner M, Chaichoompu K, et al. Benchmarking atlas-level data integration in single-cell genomics. Nature Methods. 2022;19:41–50.
  4. Armingol E, Officer A, Harismendy O, et al. Deciphering cell–cell interactions and communication from gene expression. Nature Reviews Genetics. 2021;22:71–88.
  5. Kang J, Lee JH, Cha H, et al. Systematic dissection of tumor-normal single-cell ecosystems across a thousand tumors of 30 cancer types. Nature Communications. 2024;15:4067.
For research use only. Not for use in diagnostic procedures. Data availability and analysis scope are confirmed after technical and scientific review.