Journal Club: When Histology Meets Transcriptomics — Multimodal Inference for Spatial Biology

Introduction
Spatial transcriptomics (ST) has changed how we interrogate tissue biology. By coupling gene expression measurements with spatial location, platforms such as 10x Visium, Slide-seq, and Stereo-seq let researchers move past dissociated cell suspensions and ask where cell types, signaling programs, and disease states sit within intact tissue. The technology comes with familiar caveats, though: capture spots often span multiple cells, gene expression is sparse and noisy, batch effects are pervasive, and matched histology images are not always available — or reliable proxies for transcription.
This journal club covers three recent papers that tackle these problems from complementary angles. Together they show a field-wide shift away from unimodal, deterministic models and toward multimodal, spatially aware, and uncertainty-aware approaches for tissue analysis. S²potAE (Chen et al., Briefings in Bioinformatics, 2026) integrates transcriptomics, spatial coordinates, and histology features to deconvolve cell-type composition at every spot. Li et al. (Genome Research, 2026) then apply single-cell and spatial transcriptomics to characterize intra- and intertumor heterogeneity in ovarian high-grade serous carcinoma (HGSC), revealing subtype-specific spatial programs. Finally, Stem (Zhu et al., ICLR 2025) uses a conditional diffusion model to infer spatially resolved gene expression directly from H&E images.
What unites the three is a simple recognition: morphology, location, and molecular state are deeply entangled, and exploiting that entanglement computationally yields richer insight than any single modality on its own.
Paper 1: S²potAE — A Multimodal Autoencoder for Spot Deconvolution
Key Findings
- S²potAE delivers state-of-the-art spot deconvolution across simulated and real datasets, including human breast cancer, mouse brain anterior, and human dorsolateral prefrontal cortex (DLPFC).
- The framework recovers nuanced cell-type distributions that reflect tissue architecture rather than spot-level averages, and it accurately identifies tumor boundaries.
- An auxiliary pathological classification task ties molecular composition to histologically defined phenotypes and adds interpretability.
- The approach holds up under batch effects and granularity mismatches between single-cell references and ST spots — two persistent pain points.
Methodology & Study Design
S²potAE treats each ST spot as a multimodal observation: a transcriptomic vector, a 2D coordinate, and a paired histology patch. It encodes these through (i) a graph-based spatial encoder that propagates information across neighboring spots to capture tissue-scale context, and (ii) a perceptual image encoder that extracts morphology-aware embeddings from histological crops. A multilevel feature aggregation strategy fuses the two streams before a decoder predicts cell-type proportions. The auxiliary classification head predicts tissue/pathology labels from the latent representation, encouraging the model to encode biologically meaningful features rather than reconstruction artifacts.
The authors evaluate across multiple simulated benchmarks (where ground-truth proportions are known) and several real ST datasets spanning human and mouse tissues — an important distinction, since performance on synthetic mixtures does not always carry over to real tissue.
Significance
Deconvolution underpins nearly every downstream ST analysis: cell–cell interaction inference, niche identification, and tumor microenvironment characterization all depend on knowing what is in each spot. By combining spatial and morphological evidence with expression, S²potAE reduces the ambiguity that plagues reference-based deconvolution when reference profiles are imperfect or incomplete. The auxiliary pathology head is a nice touch — it pushes the latent space toward clinically meaningful structure instead of purely numerical fit.
Paper 2: Characterizing HGSC Heterogeneity with scRNA-seq and ST
Key Findings
- The study maps intra- and intertumor heterogeneity across 2D tissue sections in ovarian HGSC and links bulk-defined molecular subtypes to spatial domains with distinct gene expression programs.
- Differential spatial patterns emerge for immune pathways and vasculature development, varying across patients and subtypes.
- Functional analysis of tumor regions points to potentially shared cellular states across molecular subtypes despite transcriptomic differences.
- Subtype-specific spatial anti-colocalization appears between spots exhibiting antigen-presentation functions and those enriched for B-cell-mediated immunity.
- Spatially aware cell–cell communication analysis picks up a molecular-subtype-specific difference in total signaling activity and heterogeneity in Midkine (MDK) signaling between differentiated subtypes.
- A practical recommendation: multiple tissue slices per patient may be necessary to comprehensively capture HGSC spatial transcriptomes.
Methodology & Study Design
Rather than introducing a new algorithm, the authors apply existing single-cell and spatial transcriptomics tools in an integrative framework. They combine scRNA-seq-derived cell-type references with ST data, perform spatial domain identification, quantify colocalization between functional gene-expression programs, and use spatially aware ligand–receptor inference to compare signaling landscapes across molecular subtypes. The emphasis is on variability — between patients (intertumor) and within a single section (intratumor) — and on linking spatial structure to clinically defined subtypes.
Significance
This is a useful reminder that methodological innovation is not the only way to advance spatial biology. Rigorous biological characterization with existing tools can surface patterns — subtype-specific immune anti-colocalization, MDK signaling differences — that motivate new hypotheses and downstream experiments. The explicit call to sample multiple slices per patient is itself a methodological contribution: it pushes back against the assumption that one section tells the whole story.
Paper 3: Stem — Diffusion-Based Gene Expression Inference from H&E
Key Findings
- Stem uses a conditional diffusion model to predict spatially resolved gene expression from H&E-stained images, achieving state-of-the-art performance on multiple tissue types and ST platforms.
- Generated profiles preserve gene variation levels similar to ground truth, suggesting the model captures underlying biological heterogeneity rather than collapsing to mean expression.
- Because the model is generative, it produces a distribution of plausible expression profiles for a given image patch, explicitly representing the inherent one-to-many mapping between morphology and transcription.
- Predictions remain biologically meaningful on held-out H&E images at test time, which hints at using the approach on archived pathology slides without paired ST.
Methodology & Study Design
Stem frames image-to-expression inference as a conditional generation problem. A diffusion model learns to reverse a noising procedure, conditioned on image features (often leveraging pathology foundation-model representations). At inference, it iteratively denoises a random latent vector into a gene-expression vector conditioned on the input image patch. Sampling multiple times yields multiple plausible outputs for the same image — a deliberate departure from deterministic regression, which returns a single point estimate.
The authors evaluate on datasets from different tissue sources and sequencing platforms, which matters because cross-platform generalization is a known weak point for image-to-expression models.
Significance
Stem is conceptually important: it treats stochasticity as a feature, not a bug. Morphologically similar tissue can be molecularly distinct (and vice versa), and a single point estimate misrepresents that uncertainty. By modeling a distribution, Stem returns predictions that better reflect biological variability. Practically, it opens the door to mining the enormous archives of H&E slides — abundant in clinics and biobanks — for molecular hypotheses when ST data are not available.
Synthesis & Discussion
A Shared Theme: Multimodal Context Wins
Read together, the three papers reinforce one message: tissue biology cannot be reconstructed from any one modality. S²potAE combines RNA + coordinates + morphology to estimate cell composition. Stem combines morphology + diffusion-based generation to estimate expression. Li et al. combine scRNA-seq + ST + spatial statistics to estimate tumor architecture. In every case, performance or insight improves when complementary signals are fused.
Different Goals, Complementary Methods
The papers differ in what they infer:
- S²potAE infers cell-type composition per spot (deconvolution).
- Stem infers gene-expression distributions per image region (generative prediction).
- Li et al. infer spatial structure and signaling patterns within tumors (biological discovery).
These are not competing approaches — they are layers. A reasonable spatial-biology workflow might use Stem-like inference to scout expression, S²potAE-like deconvolution to assign cell types, and HGSC-style spatial analyses to characterize niches and signaling.
Tensions and Caveats
- Validation depth varies. S²potAE reports accuracy gains over baselines, but independent replication on additional platforms, tissues, and patient cohorts is needed. Stem performs well within studied tissues; cross-lab, cross-scanner generalization remains an open question. Li et al.'s biological findings are correlative and call for functional follow-up.
- Uncertainty is handled unevenly. Stem explicitly models predictive uncertainty; S²potAE returns proportions as point estimates; Li et al. rely on statistical testing across patients and spots. Calibration of inferred proportions and expression — how confident should we be in any single prediction? — is underexplored.
- Morphology ≠ ground truth. Both S²potAE and Stem assume that histology contains extractable information about molecular state. That is partially true but breaks down for cell states (early activation, transient signaling) that lack strong morphological correlates.
Emerging Trends
- Multimodal fusion is now the default rather than the exception in ST analysis.
- Probabilistic and generative models (diffusion, VAEs, ensembles) are replacing single-point estimators.
- Foundation models for pathology are becoming standard feature extractors for histology-aware ST tools.
- Spatial niche and signaling analysis (cell–cell communication, colocalization) is moving from descriptive to subtype- and patient-stratified.
- Cross-platform and cross-cohort generalization is increasingly treated as a first-class evaluation criterion.
Open Questions
- How do these methods perform under realistic batch effects — varying staining protocols, scanners, fixation methods?
- Can deconvolution or generative inference substitute for true ST when clinical decisions are at stake?
- How should uncertainty from generative models be propagated into downstream analyses (differential expression, niche detection, survival models)?
- What is the minimum number of slices or spots per patient needed for stable spatial inference?
- How transferable are models trained on one tissue or disease to another?
For Your Lab Meeting
- S²potAE uses an auxiliary pathology classification head to "guide" its latent space. What are the trade-offs of adding task-specific supervision to a representation model? Could this bias the representation toward features useful for the head but irrelevant for deconvolution?
- Stem models a distribution of plausible expression profiles rather than a single prediction. In what downstream analyses would this distributional output be most valuable? Where might it cause problems (e.g., differential expression testing)?
- Li et al. recommend multiple slices per patient. Given cost and tissue availability constraints, how would you design a study to determine the minimum number of slices needed for stable spatial characterization in HGSC or another cancer type?
- All three papers implicitly assume that morphology carries extractable molecular information. For which biological questions is this assumption most likely to break down, and how could we detect such cases systematically?
- Foundation models for pathology (e.g., UNI, CONCH, CTransPath) are increasingly used as feature extractors in ST tools. How concerned should we be about subtle domain shifts when these models are applied to slides from different institutions, stains, or scanners?
Key Terms
- Spatial transcriptomics (ST): A family of technologies that measure gene expression at defined locations within a tissue section, preserving spatial context lost in dissociated scRNA-seq.
- Spot deconvolution: The computational task of estimating the proportion of cell types contributing to a single ST capture spot, which typically contains RNA from multiple cells.
- Conditional diffusion model: A generative model that learns to reverse a gradual noising process, conditioned on an input (e.g., an image), enabling sampling of plausible outputs for a given context.
- High-grade serous carcinoma (HGSC): The most common and aggressive subtype of ovarian cancer, characterized by extensive genomic instability and a complex tumor microenvironment.
- Cell–cell communication inference: Computational methods that infer ligand–receptor interactions and signaling activity from transcriptomic data, often incorporating spatial proximity to prioritize likely interactions.
References
Chen T, Xue W, Zhang YF, Luo Y, Liu C, Shen W, et al. (2026). S2potAE: multimodal spatial spot autoencoder integrating image and transcriptomic features for deconvolution. Briefings in Bioinformatics. https://doi.org/10.1093/bib/bbag020
Li W, Grieshober L, Gertz J, Ivich A, Doherty JA, Greene CS, et al. (2026). Characterizing intra- and intertumor heterogeneity in ovarian high-grade serous carcinoma subtypes using single-cell and spatial transcriptomics. Genome Research. https://doi.org/10.1101/gr.281433.125
Zhu S, Zhu Y, Tao M, & Qiu P (2025). Diffusion Generative Modeling for Spatially Resolved Gene Expression Inference from Histology Images. International Conference on Learning Representations. https://doi.org/10.48550/arXiv.2501.15598