How to Design a Spatial Transcriptomics Project: From Biological Question to Experimental Strategy
Figure 1: A stepwise decision framework for designing a spatial transcriptomics project — from defining the question to preparing the analysis plan.
Spatial transcriptomics is not a single technique — it is a category of technologies spanning spot-level capture arrays, imaging-based single-molecule detection, and subcellular-resolution sequencing. The hardest part of designing a spatial transcriptomics project is not choosing a platform. It is defining a question that genuinely needs spatial information, selecting samples and replicates that can answer it, and planning the analysis before the first tissue section is cut. This article provides a decision framework for researchers and project teams moving from "we should add spatial" to a concrete, budgeted experimental strategy, drawing on practical guidance from over 1,000 spatial samples and published power-analysis tools.
When Spatial Resolution Matters
A spatial transcriptomics project should begin with an uncomfortable question: does my hypothesis actually need spatial information?
A 2026 practical guide based on over 1,000 spatial samples identified this as the single most common failure point in study design. Researchers often skip directly to platform selection without first confirming that tissue architecture is part of the signal they are trying to detect.
The distinction is practical, not philosophical. If your question is "which genes change between treated and untreated groups," bulk RNA-seq or single-cell RNA-seq will typically give you more statistical power per dollar. Spatial information adds value when the question involves one of the following:
- Localization-dependent biology. Where in the tissue a transcript is expressed determines its functional interpretation — tumor margin vs. core, peri-necrotic zones, or cortical layer-specific expression.
- Cell-cell interactions anchored to tissue structure. Ligand-receptor co-expression that varies by anatomical compartment cannot be resolved by dissociated single-cell data alone.
- Tissue domains not defined by cell type alone. Necrotic zones, fibrotic scars, and tertiary lymphoid structures are spatial constructs — they are defined by their organization, not by a unique cell population.
- Gradients across tissue axes. Luminal-to-basal, portal-to-central, or pia-to-white-matter gradients require spatial coordinates to detect and quantify.
If your project does not fall into one of these categories, a matched bulk RNA-seq or single-cell RNA-seq approach may be the stronger choice. If it does, the next step is to pick the right spatial readout.
| Research Question Type | Best Approach | Why |
|---|---|---|
| Differential expression between groups (no spatial component) | Bulk RNA-seq or scRNA-seq | Higher power per dollar; spatial adds cost without analytical benefit |
| Gene expression gradients across a tissue axis | Spatial transcriptomics | Gradients cannot be recovered from dissociated cells |
| Cell-type co-localization in specific anatomical compartments | Spatial transcriptomics | Tissue architecture is the signal |
| Large-scale atlas of a whole organ or embryo | Stereo-seq or Visium HD | cm-scale capture with near-cellular resolution |
| Immune infiltration at the tumor-stroma interface | Imaging-based ST (Xenium, CosMx) | Single-cell resolution of cell-cell adjacencies |
Platforms, Resolution, and Coverage
Once a spatial question is confirmed, platform selection becomes a three-way trade-off — and no platform optimizes all three axes simultaneously.
Spatial resolution, gene coverage, and sample compatibility define the decision space. Sequencing-based platforms (Visium, Visium HD, Stereo-seq) offer broad or whole-transcriptome coverage but vary in resolution from 55-um spots to subcellular bins. Imaging-based platforms (Xenium, CosMx SMI, MERSCOPE) deliver single-cell or subcellular resolution but are limited to targeted gene panels of hundreds to a few thousand genes.
A 2025 BMC Genomics review categorized seven major commercial platforms into three groups, and the categories remain the most practical starting point for selection.
| Platform Group | Resolution | Gene Coverage | FFPE-Compatible | Best For |
|---|---|---|---|---|
| 10x Visium | 55 um spots (multi-cell) | Whole-transcriptome | Yes | Discovery on standard tissue structures |
| Visium HD, Stereo-seq | Near-single-cell (2–10 um bins) | Whole-transcriptome | Yes (Visium HD) / Fresh-frozen (Stereo-seq) | Discovery with fine spatial detail; organ-scale atlases |
| Xenium, CosMx SMI, MERSCOPE | Subcellular (imaging-based) | Targeted panels (100–6,000 genes) | Yes | Cell-cell interactions, immune niches, focused validation |
The resolution-coverage trade-off has a practical consequence that is easy to overlook: a whole-transcriptome method can reveal patterns you did not expect, while a targeted panel can only confirm or reject hypotheses you already have. If your gene list is incomplete — or the biology is incompletely understood — a discovery-grade approach may save you from missing the most interesting signal.
For researchers ready to work with a specific platform, CD Genomics offers spatial transcriptomics services across multiple technologies, including 10x Visium HD for high-resolution whole-transcriptome profiling and Stereo-seq for subcellular-resolution organ-scale mapping. For a deeper comparison of platform trade-offs, see the platform selection guide.
Samples, Sections, and Statistical Power
Sample strategy is where most spatial transcriptomics projects succeed or fail — and it is the part of the design process that receives the least structured attention.
Fresh-frozen vs. FFPE. Fresh-frozen tissue generally preserves higher RNA integrity and supports a wider range of platforms. FFPE samples are far more abundant in biobanks and pathology archives but introduce RNA degradation that demands higher sequencing depth. A 2026 Trends in Biotechnology review reported that FFPE Visium experiments routinely require 100,000–120,000 reads per spot — roughly four times the historical 25,000-read recommendation — to achieve comparable gene detection. The FF vs. FFPE decision guide covers tissue handling considerations in more detail.
Biological replicates vs. regions of interest. Statistical power in spatial transcriptomics is driven primarily by the number of independent biological samples — not by the number of sections or regions of interest per sample. Ryaboshapkina and Azzu (2023) demonstrated this quantitatively: in a GeoMx DSP study of liver fibrosis, power to detect expression differences between patient groups increased with patient number but showed diminishing returns from additional ROIs within the same patient. Ten patients per group with two ROIs each provided approximately 80% power for the primary endpoint.
Baker et al. (2023) generalized this principle by introducing in silico tissue (IST) generation — a simulation framework that lets researchers model statistical power for cell-type detection and differential cell-cell adjacency testing before committing to a full experiment. The method works with pilot data from any spatial platform and can answer design questions such as "how many fields of view do I need to detect a 1.5-fold change in a rare cell population?"
Within-sample replication still matters for QC. Two non-adjacent sections per sample provide insurance: if one section fails QC due to folding, tearing, or low RNA recovery, the second section preserves the sample. This is standard practice in core facilities that process high volumes of spatial samples.
| Design Factor | Recommendation | Rationale |
|---|---|---|
| Biological replicates per group | 5 or more for discovery; 10 or more for formal statistical testing | Patient-to-patient variation is the dominant source of noise |
| Sections per sample | 2 or more (non-adjacent) | QC backup; within-sample consistency check |
| ROIs per section | Match to anatomical compartments; 2 or more per compartment | Capture intra-regional heterogeneity |
| Pilot experiment | Run on 2–3 samples before scaling | Calibrate permeabilization, sequencing depth, and expected gene recovery |
| Power analysis | Use IST simulation (Baker et al.) or platform-specific tools | Avoid underpowered designs |
Figure 2: Sample and replicate strategy — balancing biological replicates, within-sample sections, and ROI selection for statistical power.
Plan Your Analysis First
A spatial transcriptomics dataset can reach hundreds of gigabytes. The time to decide how it will be analyzed is before it exists — not when a bioinformatician opens the delivery folder.
Analysis planning for a spatial project should address five questions at minimum:
- What defines a positive signal? For spatially variable genes, pre-specify which method (SpatialDE, SPARK-X, Moran's I) and what significance threshold you will use. Changing methods post hoc can change which genes are called significant.
- How will spatial domains be defined? If your hypothesis involves tissue compartments, define how those compartments will be identified — unsupervised clustering, histology-guided annotation, or reference-based mapping. Each approach has different assumptions.
- What is your cell-type annotation strategy? If using spot-level data, will you deconvolve spots with a single-cell reference, or annotate at the spot level? If the latter, how will you handle spot mixing in dense regions?
- What orthogonal validation is planned? At minimum, confirm key spatial domain boundaries against histology (H&E or IF). For high-stakes claims — a new cell state, a drug-targetable interaction — plan for RNAscope or immunofluorescence validation of the spatial pattern.
- Who will run the analysis, and on what infrastructure? A Stereo-seq bin20 full-chip dataset requires HPC or cloud compute. Underestimating storage and processing needs is one of the most common budget overruns in spatial projects.
The spatial transcriptomics data analysis guide covers workflows, tools, and QC expectations in detail. For teams that prefer to hand off computational work, CD Genomics spatial transcriptomics data analysis provides pipeline-based processing with customizable deliverables.
Adding scRNA-seq or Multi-Omics Data
Not every spatial project needs an additional modality. But the ones that do benefit from scRNA-seq or multi-omics integration tend to fall into two clear scenarios.
Scenario 1: Spot-level deconvolution. Visium and Visium HD spots capture mRNA from multiple cells. If your question requires cell-type-specific resolution within spots, a matched scRNA-seq or snRNA-seq reference dataset enables deconvolution — estimating the proportion of each cell type at each spatial location. This is the most common reason to add single-cell data to a spatial project. The requirement is that the single-cell reference and the spatial samples come from the same or closely matched tissue context.
Scenario 2: Mechanism beyond transcript abundance. Spatial transcriptomics measures RNA. If your hypothesis involves chromatin accessibility, protein localization, or metabolite distribution, adding a spatial epigenomics or spatial proteomics layer can connect transcript patterns to regulatory or functional readouts. This adds substantial cost and analytical complexity — the decision should be driven by a specific biological question, not by "more data is better."
Signs that you should keep it simple:
- Your spatial data alone resolves clear tissue domains with interpretable marker gene patterns.
- Deconvolution is not needed because your cell populations are anatomically segregated.
- Adding a second modality would stretch the budget without a hypothesis for what the additional data would answer.
For projects that genuinely need integrated readouts, the scRNA-seq and spatial transcriptomics integration guide covers deconvolution tools, alignment methods, and practical validation strategies.
Figure 3: Analysis planning workflow — key decision points to resolve before data generation, from QC thresholds through spatial domain definition and orthogonal validation.
Design Traps That Undermine Projects
Some design errors are recoverable. The ones listed here usually are not.
- Choosing a platform before defining the question. This is the most frequent mistake in spatial study design. A platform that excels at whole-transcriptome discovery may be the wrong tool for quantifying pre-specified cell-cell interactions — and vice versa.
- Treating one section as representative of a whole tissue. Heterogeneous tissues — tumors, inflammatory lesions, developing organs — vary dramatically across sections. A single 10-um section from a 5-mm biopsy represents roughly 0.2% of the tissue volume. Multi-region sampling with histological anchoring is the minimum standard.
- Confusing technical replicates with biological replicates. Two sections from the same tissue block are technical replicates. They tell you about section-to-section variability, not patient-to-patient or condition-to-condition variability. Only biological replicates provide the degrees of freedom needed for statistical inference.
- Skipping the pilot. Running a full cohort without pilot data is gambling with the budget. A pilot of 2–3 samples lets you calibrate permeabilization conditions, verify expected gene and molecular barcode recovery, estimate needed sequencing depth, and test that your analysis pipeline produces interpretable results — all before committing to 20+ samples.
- Deferring analysis planning to post-sequencing. The time to discover that your computational infrastructure cannot handle the data, or that your annotation strategy requires a reference dataset you do not have, is before tissue hits the cryostat. Retrospective analysis decisions undermine reproducibility and can make the difference between a publishable finding and an expensive data graveyard.
Planning a spatial transcriptomics project requires careful consideration of biological goals, sample availability, platform selection, and downstream analysis strategy. If you are unsure which workflow best fits your research objectives, CD Genomics experts can help evaluate your project requirements and recommend an appropriate spatial transcriptomics strategy.
FAQ
Q: How do I know whether my project needs spatial transcriptomics or whether single-cell RNA-seq is sufficient?
A: If your central question involves where transcripts or cell types are located relative to tissue structures — tumor margins, necrotic zones, cortical layers, fibrotic scars — spatial transcriptomics adds information that dissociated single-cell methods cannot recover. If your question is "which genes change between conditions" without a spatial component, scRNA-seq or bulk RNA-seq typically provides more statistical power at lower cost. When in doubt, ask whether tissue architecture appears in your hypothesis statement. If it does not, spatial data may be unnecessary.
Q: How many biological replicates do I need for a spatial transcriptomics study?
A: For discovery-phase projects using whole-transcriptome methods, five biological replicates per condition is a practical minimum. For formal statistical comparisons — differential expression between groups, differential cell-cell adjacency testing — ten or more replicates per group is recommended. The key finding from published power analyses is that adding more patients increases statistical power far more than adding more ROIs or sections within the same patient. A pilot experiment combined with in silico tissue simulation can provide project-specific power estimates.
Q: Should I use fresh-frozen or FFPE tissue for my spatial transcriptomics project?
A: Fresh-frozen tissue generally yields higher RNA quality and is compatible with a wider range of platforms, including Stereo-seq. FFPE tissue is more abundant — especially in clinical archives — and is supported by Visium, Visium HD, and most imaging-based platforms, but it requires deeper sequencing (typically 100,000–120,000 reads per spot for Visium FFPE) to achieve comparable gene detection sensitivity. The choice should be driven by sample availability and the platform you plan to use, not by a universal preference for one preservation method.
Q: When should I add scRNA-seq data to my spatial transcriptomics project?
A: Add scRNA-seq or snRNA-seq when you need cell-type-level resolution from spot-based spatial data (via deconvolution), or when you need to validate that cell states identified in spatial data correspond to discrete transcriptomic clusters. If your cell populations are anatomically segregated — for example, neurons in clearly defined cortical layers — deconvolution may be unnecessary and the single-cell budget can be redirected to additional spatial samples.
Q: What is the single most important thing to do before starting a spatial transcriptomics project?
A: Run a pilot experiment on 2–3 representative samples. A pilot calibrates permeabilization, verifies RNA quality and expected gene recovery, estimates the sequencing depth you will actually need, and confirms that your analysis pipeline produces interpretable results. The cost of a pilot is a small fraction of a failed full-scale experiment, and the information it provides — about tissue handling, platform behavior, and analytical feasibility — is project-specific and irreplaceable.
References
- Grases D, Porta-Pardo E. A practical guide to spatial transcriptomics: lessons from over 1000 samples. Trends in Biotechnology. 2026;44(5):1230-1242. doi:10.1016/j.tibtech.2025.08.020
- Baker EAG, Schapiro D, Dumitrascu B, Vickovic S, Regev A. In silico tissue generation and power analysis for spatial omics. Nature Methods. 2023;20(3):424-431. doi:10.1038/s41592-023-01766-6
- Righelli D, Sottosanti A, Risso D. Designing spatial transcriptomic experiments. Nature Methods. 2023;20(3):355-356. doi:10.1038/s41592-023-01801-6
- Lim HJ, Wang Y, Buzdin A, Li X. A practical guide for choosing an optimal spatial transcriptomics technology from seven major commercially available options. BMC Genomics. 2025;26:47. doi:10.1186/s12864-025-11235-3
- Ryaboshapkina M, Azzu V. Sample size calculation for a NanoString GeoMx spatial transcriptomics experiment to study predictors of fibrosis progression in non-alcoholic fatty liver disease. Scientific Reports. 2023;13:8943. doi:10.1038/s41598-023-36187-0
The information provided in this article is for research use only and is not intended for use in diagnostic or therapeutic procedures. CD Genomics provides sequencing and bioinformatics services for research purposes. Researchers should consult the appropriate regulatory guidelines for their specific applications.