Journal Club: Spatial GRNs, Foundation-Model Fine-Tuning, and Deconvolution

· 11 min read · Paper Feed

Journal Club: Spatial GRNs, Foundation-Model Fine-Tuning, and Deconvolution

Conceptual illustration: Three New Frontiers in Spatial Transcriptomics—From Regulatory Networks to Foundation Models to Deconvolution

Introduction

Spatial transcriptomics (ST) has reshaped how we interrogate tissue biology, moving us beyond dissociated cell atlases toward measurements in which gene expression is anchored to physical coordinates. The technology now spans platforms as varied as Visium, seqFISH, Xenium, and MERFISH, generating data at resolutions from multi-cellular spots to near-single-cell subcellular maps. Each measurement is only as useful as the analytical framework we bring to it, and the community is currently wrestling with three intertwined questions: How do we model spatial context explicitly? How do we adapt large models without owning a GPU cluster? And how do we infer cell-type composition when our 'spots' are mixtures?

This journal club synthesizes three recent papers that, taken together, sketch the current state of the field. Li et al. introduce SVGRN, a deep generative framework for inferring spatially varying gene regulatory networks at spot or cell resolution. Zou and Lei present SpatialPEFT, a parameter-efficient fine-tuning toolkit that lets researchers adapt ST foundation models of up to 1.4 billion parameters on a single 16 GB consumer GPU. Ding, Yu, and Ming describe AddaGCN, a graph convolutional network with adversarial domain adaptation for ST deconvolution. Each tackles a different link in the spatial-analysis pipeline, and reading them side by side reveals a community converging on graph-based, deep, and context-aware methods, while still grappling with interpretability, validation, and accessibility.

Paper 1: SVGRN — Spot-Specific Gene Regulatory Networks That Actually Respect Position

Li, Chen, Lu, Tsai & Wang (2026), Bioinformatics Advances — DOI: 10.1093/bioadv/vbag230

Key Findings

  • Spatially varying, not spatially uniform, regulation. SVGRN infers a separate GRN for each spot or cell, conditioned on that location and its neighborhood, rather than collapsing tissue into a single consensus network.
  • Consistent gains across simulated and real data. The framework outperforms existing GRN inference methods on simulated benchmarks and is demonstrated on seqFISH mouse embryo data, Visium human cutaneous squamous cell carcinoma (cSCC), and fallopian tube samples, capturing regulatory programs associated with developmental boundaries, tumor progression, and tissue compartmentalization.
  • Nonlinear and context dependent. By combining a conditional variational autoencoder (CVAE) with a structural equation model (SEM), SVGRN learns nonlinear regulatory programs rather than simple pairwise correlations.
  • Unsupervised, prior-aware. Candidate regulatory edges from prior knowledge (e.g., TF–target databases) are integrated as inputs, allowing the model to refine them in a data-driven way without requiring curated ground truth.

Methodology & Study Design

The authors frame gene regulation as a structural equation: expression of a target gene is a function of its regulators, its spatial location, and the local neighborhood. They implement this with a CVAE that takes (i) target gene expression, (ii) candidate regulator expression, (iii) spot coordinates, and (iv) neighborhood features as conditioning inputs. The latent variables encode location-specific regulatory structure, and the decoder reconstructs gene expression under the implied GRN. The model is trained unsupervised, so it can be deployed in tissues where perturbation data are unavailable. They validate on three real datasets: a seqFISH mouse embryo (targeting a few hundred genes in a known developmental context), a Visium human cSCC sample for tumor–stroma contrasts, and a Visium fallopian tube sample for epithelial organization.

Significance

Most GRN tools, such as GRNBoost2, PIDC or SCENIC+, infer one network for a whole dataset or cell population and never look at where the cells sit in the tissue. SVGRN's central insight is that the same cell type can adopt different regulatory states in different tissue niches, and that this is detectable from expression plus spatial location alone. For tumor biology, this is appealing: SVGRN can in principle identify regulators active only at the invasive front, only in hypoxic cores, or only in immune-infiltrated regions. The main caveats are familiar: edges are inferred associations, candidate regulators must be supplied, and performance will degrade with very sparse or noisy ST data.

Paper 2: SpatialPEFT — Fine-Tuning 1.4-Billion-Parameter ST Models on a 16 GB GPU

Zou & Lei (2026), Bioinformatics — DOI: 10.1093/bioinformatics/btag503

Key Findings

  • 87% lower peak VRAM. On Geneformer-316M, full fine-tuning runs out of memory on a 16 GB card; LoRA plus gradient checkpointing brings peak VRAM down to 2.15 GB (an 87.2% reduction, per the paper). The largest supported model, the 1.4-billion-parameter UCE, fine-tunes in 9.48 GB.
  • Large accuracy gains with <1% trainable parameters. On a Xenium FFPE breast cancer dataset (about 576,000 cells, 11 classes), cell-type annotation Macro F1 rose from 0.7104 (zero-shot) to 0.9586 while training only 0.20% of parameters.
  • Single consumer GPU is enough. Reported runs use a single RTX 4080 Super (16 GB), removing the institutional compute barrier for many labs.
  • Spatial-aware adapter matters. Generic LoRA alone loses tissue-context information; the authors' adapter module is designed to preserve spatial structure during adaptation.

Methodology & Study Design

SpatialPEFT is an engineering framework rather than a new model architecture. It wraps an existing ST foundation model and injects LoRA modules into selected linear layers, enabling low-rank updates to attention and feed-forward blocks. Gradient checkpointing recomputes activations during backward pass to save memory. The spatial-aware adapter is a lightweight module that consumes neighborhood information alongside token embeddings, ensuring that the fine-tuned model remains sensitive to tissue geometry. The toolkit is implemented in Python, MIT-licensed, and shipped with documentation and tutorials. They report results on Xenium breast cancer (cell-type annotation) and the 12-slice DLPFC Visium dataset (spatial domain identification).

Significance

Foundation models for ST, such as scGPT-spatial, Geneformer derivatives, and platform-specific encoders, are becoming reusable biological infrastructure, but full fine-tuning is often prohibitively expensive. SpatialPEFT slots into the broader PEFT (parameter-efficient fine-tuning) trend that has already transformed NLP and is now spreading to genomics. For experimentalists, the practical implication is meaningful: you can take a pretrained ST foundation model, fine-tune it on your own tissue or disease cohort on a workstation, and ship it without renting cloud GPUs. The flip side is that evaluation is still task-specific, and catastrophic forgetting remains a risk; the paper's reported F1 gains are impressive but were measured on annotation tasks where labeled data are available, not on open-ended representation quality.

Paper 3: AddaGCN — Graph Convolutional Deconvolution with Adversarial Domain Adaptation

Ding, Yu & Ming (2026), PLOS Computational Biology — DOI: 10.1371/journal.pcbi.1014609

Key Findings

  • Robust cell-type deconvolution across platforms. AddaGCN reports superior performance on real ST data from multiple technologies (including 10x Visium and Stereo-seq) compared with existing deconvolution methods.
  • Domain adaptation reduces reference–query mismatch. An adversarial discriminative component aligns ST data with a single-cell reference, mitigating batch effects that would otherwise distort cell-type proportion estimates.
  • Spatiotemporal and tumor-microenvironment applications. The authors demonstrate AddaGCN's ability to track developmental changes and characterize tumor microenvironments, including tumor–immune interfaces.
  • Graph structure is load-bearing. Spatial neighbors are encoded as a graph, and a graph convolutional network (GCN) propagates information across adjacent spots before deconvolution.

Methodology & Study Design

Deconvolution in ST aims to estimate the proportion of each cell type at every spot, given a high-resolution single-cell reference. AddaGCN builds a spatial graph over spots, feeds expression through a GCN to produce spatially smoothed features, and then runs an adversarial discriminative domain adaptation step, a GAN-style objective where a domain classifier tries to distinguish "reference" features from "ST" features, and the GCN learns to fool it. The output is a per-spot, per-cell-type proportion vector. Compared with methods like Stereoscope, cell2location, SPOTlight, and DSTG, the authors report improved accuracy and robustness, especially when reference and query come from different platforms or donors.

Significance

Deconvolution is often the first analytical step in an ST pipeline, and its errors propagate everywhere downstream. AddaGCN addresses the most common failure mode: reference–query distribution shift, which arises because ST and scRNA-seq differ in capture chemistry, gene coverage, and sampling biases. The goal overlaps with batch-integration tools such as scVI or Seurat's anchor-based integration, which also put reference and query into a shared space; AddaGCN gets there with an adversarial objective and aims it at the ST–scRNA bridge specifically. The paper's tumor microenvironment results are particularly useful because reliable cell-type proportions are exactly what you need before asking the kinds of regulatory and spatial-niche questions raised in SVGRN.

Synthesis & Discussion

Read individually, each paper solves a different problem. Read together, they reveal a coherent picture of where spatial transcriptomics is heading.

Shared methodological DNA. All three methods use graph-based spatial structure as a first-class citizen rather than a post-hoc visualization. SVGRN conditions its CVAE on spatial coordinates and neighborhoods; SpatialPEFT injects a spatial-aware adapter; AddaGCN builds a GCN over the spot graph. Spatial transcriptomics is no longer "scRNA-seq plus x/y"; the geometry is now part of the model architecture, not an afterthought.

Complementary rather than competing. The three papers are designed to stack. AddaGCN gives you cell-type proportions; SVGRN then operates on those compositions (or on the expression itself) to give you regulatory mechanisms; SpatialPEFT provides the computational substrate that could host either downstream task via a foundation model. A natural pipeline emerges: deconvolve with AddaGCN, infer spatially varying GRNs with SVGRN, and if you need a custom foundation model for your tissue, fine-tune it with SpatialPEFT. None of the papers claim to replace the others, and their evaluation benchmarks barely overlap, which is both a strength (they're solving real, narrow problems well) and a weakness (we still lack head-to-head comparisons on shared tasks).

Where the field still has gaps.

  • Causality vs. correlation. SVGRN's edges are inferred associations, even when built on a structural-equation scaffold. None of the three papers proposes perturbation-aware validation as a default.
  • Uncertainty quantification. None of the three papers provides calibrated confidence intervals on its outputs, a known weakness of most deep-learning ST tools.
  • Resolution heterogeneity. SVGRN is evaluated on seqFISH (subcellular) and Visium (multi-cellular) data; AddaGCN on multi-cellular platforms; SpatialPEFT on Xenium (single-cell) and Visium DLPFC (multi-cellular). It's unclear how each behaves on the platforms it wasn't designed for.
  • Standardized benchmarks. The field still lacks an ST analogue of the scRNA-seq benchmark suites (e.g., scIB), making cross-paper comparisons difficult.
  • Foundation models for regulation. It's an open question whether the same pretrained ST encoder that SpatialPEFT adapts for annotation could also be adapted to output spatially varying GRNs, blurring the line between papers 1 and 2.

For Your Lab Meeting

  1. Methodological: SVGRN combines a CVAE with a structural equation model, but the SEM imposes directional structure. Should we view SVGRN's outputs as causal hypotheses, or strictly as associations with directional prior? What experiments would distinguish the two?
  2. Generalizability: All three papers demonstrate on one or two tissues each. How would you design a benchmark to test SVGRN, SpatialPEFT, and AddaGCN on a single shared dataset (e.g., a public Xenium tumor cohort)?
  3. Foundation models: SpatialPEFT adapts encoders for annotation. Could a similar LoRA approach be used to adapt a foundation model to output spatially varying GRNs? What would be lost compared with SVGRN's explicit SEM?
  4. Deconvolution as upstream input: If we run AddaGCN first, then feed its cell-type proportions into SVGRN, do we introduce circularity? Or does decomposing the signal into known cell types actually help the regulatory inference?
  5. Validation: What is the minimal experimental validation (e.g., spatial FISH, Perturb-seq, ATAC-seq) that would give you confidence in a spatially varying GRN edge inferred by SVGRN in a tissue like cSCC or fallopian tube?

Key Terms

  • Structural Equation Model (SEM): A statistical framework that models observed variables as a system of linear (or nonlinear) equations, allowing directional dependencies between variables rather than mere correlations.
  • Conditional Variational Autoencoder (CVAE): A generative neural network that learns a probabilistic mapping from inputs to a latent space, conditioned on auxiliary variables (here: spatial location and neighborhood features).
  • Parameter-Efficient Fine-Tuning (PEFT) / LoRA: A family of techniques for adapting large pretrained models by training only a small subset of parameters (e.g., low-rank matrices inserted into frozen layers), drastically reducing compute and memory.
  • Adversarial Domain Adaptation: A training strategy in which a feature extractor is optimized to fool a domain classifier, forcing it to produce representations that are indistinguishable across source (e.g., scRNA-seq) and target (e.g., ST) domains.
  • Deconvolution (in ST): The computational problem of estimating the proportion of each cell type contributing to a multi-cellular spatial spot, given a high-resolution single-cell reference.

References

  1. Li Y, Chen J, Lu T, Tsai NP, & Wang H (2026). Spatially varying gene regulation network inference from spatial transcriptomics. Bioinformatics advances. https://doi.org/10.1093/bioadv/vbag230

  2. Zou X & Lei X (2026). SpatialPEFT: a parameter-efficient fine-tuning framework for spatial transcriptomics foundation models. Bioinformatics (Oxford, England). https://doi.org/10.1093/bioinformatics/btag503

  3. Ding S, Yu Z, & Ming J (2026). AddaGCN: Spatial transcriptomics deconvolution using graph convolutional networks with adversarial discriminative domain adaptation. PLoS computational biology. https://doi.org/10.1371/journal.pcbi.1014609

Spatial Transcriptomics Spatial Omics Single-Cell