Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

There is no reliable universal number of samples for a spatial molecular study. The right sample size depends on what you want to detect, how much biological variation exists between donors or animals, how tissue is organized, and how the samples will be measured and analyzed. For a comparison intended to generalize across people or animals, independent donors or animals—not the number of cells, spots, or fields—are the key biological replication units.

A defensible estimate starts with one primary endpoint and a minimum meaningful effect, then uses pilot data or simulations that reflect the planned assay, sampling geometry, and analysis.

What does “sample size” mean in a spatial study?

Spatial molecular studies have several nested measurement levels. A study may include donors or animals, tissue specimens, sections, slides, regions of interest (ROIs) or fields of view (FOVs), and then cells, spots, or bins within each region. These counts answer different questions and should not be combined into a single sample-size number.

Biological replicates support generalization

For a condition comparison intended to generalize across people or animals, the donor or animal is usually the relevant biological unit. Bioconductor’s OSTA design guidance distinguishes this biological unit from the experimental unit—the smallest unit independently assigned to a condition—and the observational unit where a measurement is made.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For Visium, Visium HD or Stereo-seq, and CosMx or Xenium, measurements may be made at the spot, bin, or segmented-cell level. Those observations do not automatically become independent experimental replicates. Treating millions of cells from a few donors as millions of independent samples can produce pseudoreplication and overstate how much independent evidence the study contains.

Technical measurements improve characterization, not biological replication

More cells or spots, additional ROIs, repeated sections, and repeat runs can improve precision or spatial coverage for a specimen. They do not, by themselves, add independent donors or animals to a group comparison. Keep biological-replicate counts separate from sections, slides, ROIs/FOVs, and measurement-unit counts when planning and reporting the study.

Define the endpoint before estimating a number

“Power the spatial experiment” is not a sufficiently specific objective. Different endpoints have different data structures and sampling needs. Choose one primary biological endpoint, one primary contrast, and the smallest effect that would be meaningful to detect before settling on a sample-size target.

  • Differential expression: Which genes are expected to change, by how much, and under what spatially adjusted analysis?
  • Cell-type detection: Is the goal to detect a cell type, including one that may be rare?
  • Cell-cell adjacency: Is the goal to detect enrichment of a defined pair of cell types in neighboring positions?
  • Tissue or cohort organization: Is the goal to compare spatial organization across tissues, donors, or conditions?

These questions cannot safely share one sample-size calculation. For example, adding cells may help characterize a rare feature within a specimen, while adding independent donors may be more important for estimating how consistently a condition-level effect appears across a population.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build a study-specific estimate

1. Identify units, assignment, and randomization

Map the study’s hierarchy: which donors or animals are included, what is independently assigned to a condition, which tissues or specimens come from each unit, and where measurements occur. Separate biological units from technical repeats such as serial sections or repeated runs on the same specimen. Where feasible, distribute conditions across processing slides and batches rather than allowing condition to be confounded with batch.

2. Set the statistical assumptions

Conventional power calculations depend on the desired error rate, the effect size to detect, and the number of independent units. Spatial studies add further dependencies: tissue coordinates and organization, feature frequency and scale, spatial coverage, assay detection properties, and variation within and between biological units.

Use relevant pilot or public data to estimate plausible between-unit variability, feature frequency, expression or detection properties, and effect sizes. If the pilot is too small to characterize tissue structure or variability well, do not present a single precise-looking estimate as settled. Show how the result changes under a sensible range of assumptions.

3. Simulate the analysis you intend to run

Whenever possible, generate or resample data under the planned design and run the intended analysis procedure on each simulated dataset. This makes the estimate responsive to the actual endpoint, sampling plan, and model rather than to a simplified calculation that may not match the final study. Check that the simulation represents the biological units, technical measurements, tissue geometry, and sources of variation relevant to the experiment.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A calculation designed for one platform or endpoint should not be transferred to another without checking its assumptions. Simulated power is conditional on the tissue model and parameter estimates being plausible; it is not a guarantee that the real tissue will behave like the simulation.

Account for spatial coverage and tissue architecture

For imaging-based assays, the number and placement of ROIs or FOVs matter alongside the number of biological replicates. A field needs to be large and appropriately located to capture the tissue structures relevant to the endpoint. Decide whether the analysis concerns a tumor region, a brain layer, a tertiary lymphoid structure, or another feature, then plan coverage around that feature’s expected scale and distribution.

Fixed or constrained imaging areas can limit what is sampled. Tissue microarrays can increase cohort throughput, but small cores may miss within-tissue heterogeneity. More fields from one specimen can make its spatial profile more representative; they still do not substitute for additional independent specimens when the claim is about variation across donors or animals.

A 2023 in-silico tissue study illustrates why no general FOV threshold works. In its simulated spleen example, sampling more than 7.5% of the assayed tissue area—approximately 123 × 123 μm, or about 5,600 cells—was estimated to recover a particular CD4+ and CD8+ T-cell adjacency as significant with 80% probability. That result applies to the paper’s tissue, adjacency definition, and simulated setup; the authors tied the inflection point to the spatial scale of organization, not to a general recommendation for other tissues or assays.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Choose a planning method that matches the question

Available approaches address different endpoints and platforms. They are complementary, not interchangeable calculators.

Approach Best-matched question Data and assumptions Scope and caveats
PoweREST Differential-expression detection for Visium spatial transcriptomics The published framework uses nonparametric bootstrap replicates within ROIs and accounts for spatial expression, condition-associated log-fold changes, gene-detection rates, and slice replicates. Designed for Visium DEG power, not all spatial modalities or endpoints. The authors describe use when preliminary spatial data are available and an interactive application based on two cancer datasets when they are not.
spaCraft Multi-sample spatial transcriptomics planning, including a spatially adjusted differential-expression endpoint and a compositional endpoint The repository describes a cohort-level generative model learned from pilot samples and generate-recover-test Monte Carlo simulations, with spatial domains rediscovered in each replicate. The repository reports validation on 10x Visium, Visium HD, and Stereo-seq. Its README states that R 4.1.0 or later and a C++ toolchain are required, and that its methods manuscript is in preparation. Check the current version, documentation, and fit to the intended study before adopting it.
In-silico tissue framework Questions such as cell-type detection, enriched cell-cell adjacency, and tissue or cohort organization Simulates tissue and sampling to examine spatial outcomes; the framework emphasizes that tissue organization can be difficult to parameterize and relevant data may be unavailable. Results depend on the plausibility of the tissue model and sampled data. It is an illustration of endpoint-specific spatial power analysis, not evidence for a universal sample count.

For any approach, check whether it represents the relevant biological and technical units, models tissue geometry and spatial dependence, uses pilot data where appropriate, and repeats the analysis procedure planned for the real data. Also examine sensitivity to plausible tissue heterogeneity and effect-size assumptions.

Report the design so the estimate can be evaluated

State the assumptions and sampling structure, not just a headline number of samples. A useful protocol or paper should report:

  • Biological units per group and the experimental or randomization unit.
  • Sections, slides, ROIs/FOVs, and measurement units per biological unit, including their placement and spatial coverage.
  • The primary endpoint, contrast, minimum effect assumption, expected variability, power target, and type-I error target.
  • The source of pilot or other input data and whether the estimate is pilot-based or an assumption-based scenario.
  • The simulation or calculation method and the exact analysis procedure run within simulations.
  • How batch effects and multiple testing are handled, plus sensitivity to alternative plausible assumptions.
  • Limitations in the available data, tissue model, or sampling plan that could change the estimate.

If the organism or patient population, tissue context, endpoint, effect size, expected between-unit variation, platform, spatial sampling plan, or statistical model is not yet specified, the evidence does not support a defensible universal number. The next step is to define those inputs and calculate a study-specific range.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Quick Recap

SaleBestseller No. 1
SaleBestseller No. 2
SaleBestseller No. 3
SaleBestseller No. 5

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.