Skip to main content
Back to timeline
Nature CommunicationsSource publication:

SpaCEy links tissue spatial patterns to clinical outcomes with an explainable graph neural network, improving prediction and surfacing spatial markers in lung and breast cancer cohorts

Synopsis

The authors present SpaCEy (Spatial Clinical Explainability), an explainable graph neural network that builds each tissue sample into a spatial graph via Delaunay triangulation, with nodes as single-cell protein-marker abundances and no predefined cell-type or anatomical-region inputs, learns predictive embeddings with a GNN, and uses a GNNExplainer-style explainer to output edge masks that are aggregated over k-hop neighbourhoods into node importance, thereby localising contiguous outcome-associated spatial regions and key proteins; it predicted progression in a 416-patient lung adenocarcinoma cohort (accuracy 0.68, F1 0.68, AUC 0.62, versus Ali et al. 0.61/0.59/0.57 and SPACE-GM 0.55/0.52/0.

Source-provided article image: SpaCEy links spatial tissue patterns to clinical outcomes using explainable graph neural networks
Fig. 1

Fig. 1: SpaCEy model architecture and analysis workflow.

Interpretation

SpaCEy models tissue as a spatial graph whose node features are measured protein-marker abundances only; cell-type and compartment annotations do not enter graph construction, training, prediction or explanation generation and are used only post hoc for interpretation. Prior spatial methods such as SPACE-GM and Ali et al. typically take predefined cell-type labels or local clustering as inputs, whereas SpaCEy discovers emergent organisational patterns directly from continuous molecular measurements. The methodological description is explicit: Delaunay triangulation for graph construction, removal of edges longer than the 99th percentile, exclusion of images with fewer than 100 cells, and annotations retained only as metadata.

An integrated explainer outputs edge masks and aggregates k-hop neighbourhoods into node importance, localising contiguous, compact spatial regions of importance and a concise set of protein markers driving predictions. Earlier models largely limited interpretability to latent representations or cell-level importance scores, whereas SpaCEy provides both spatially contiguous regions and a marker list. Ablation in JacksonFischer showed progressively removing 1% to 80% of important nodes caused larger and monotonic performance drops than removing matched non-important nodes (e.g., 0.2175 versus 0.1289 at 80%); in synthetic tumour-stroma data, intra-tumoral CD8 nodes received higher mean importance than stromal CD8 nodes (0.0467 versus 0.0162; paired one-sided Wilcoxon p=8.29×10⁻¹⁴).

Across lung adenocarcinoma progression prediction and two breast cancer survival cohorts, SpaCEy outperformed spatial and non-spatial baselines and stratified patients across clinical subtypes. Under identical fixed splits, SpaCEy exceeded Ali et al. and SPACE-GM on accuracy, F1 and AUC in LUAD, and improved C-index over the best non-spatial baseline by 16.5% relative in JacksonFischer and 20.7% in METABRIC. LUAD used five-fold cross-validation; breast cancer used a 10-fold protocol with patient-level stratification, optimising 24 non-spatial baseline configurations, and SpaCEy showed lower standard deviations (0.061 and 0.038).

Explainer-identified spatial markers partially overlap across cohorts, suggesting transferable prognosis-related signals while retaining cohort-specific variation. Both JacksonFischer and METABRIC showed enrichment of fibronectin and β-catenin in lower-survival groups and hormone receptor-positive cells and apoptotic markers in higher-survival groups; however, fibronectin was not enriched in METABRIC HR+HER2- relative to other subtypes (mean 5.62 versus 6.34, p=0.84), leading the authors to frame it as a cohort-level lower-survival feature rather than a universally conserved prognostic biomarker. Differential abundance and cell-composition analyses across two independent breast cancer cohorts, with descriptive subtype-composition comparisons and compartment-stratified marker summaries.

Perspective

The work targets multiplex imaging single-cell spatial proteomics, where node features are protein-marker abundances, and therefore provides only a partial view of cellular states and interactions; the authors note the design is broadly applicable and can extend to other disease contexts, and they validated same-panel cross-batch transfer in a colorectal cancer CODEX cohort (35 patients, 140 tissue regions; A-to-B AUPRC 0.801±0.016 versus 0.846±0.026 under random patient stratification). Scalability analysis showed favourable scaling with sample size and marker dimensionality, with the largest setting corresponding to 1.5 billion cells per epoch and peak GPU memory around 1.06 GB, supporting feasibility for whole-slide-scale application.

Explanation quality is tied to predictive model performance, and the authors caution that explanations based on modest-performing models should be interpreted conservatively; importance scores are unsigned relevance magnitudes that do not encode direction of effect and are not causal determinants. Crop-sensitivity analysis showed predictions and embeddings were largely stable (mean Spearman ρ=0.943, cosine similarity 0.920), but ROI restriction remains a relevant limitation for both prediction and interpretation, since local crops may not capture the full tissue context of the complete ROI. Moreover, current evidence is limited to multiplex imaging spatial proteomics, and computational efficiency and explanation reliability when extending to multimodal data such as transcriptomics or metabolomics remain to be tested.

Sources