BrainSTR models dynamic brain networks with spatio-temporal contrastive learning, validated on ASD, BD, and MDD with critical phases and subnetworks consistent with prior neuroimaging findings
Synopsis
The work proposes BrainSTR, a spatio-temporal contrastive learning framework that learns state-consistent phase boundaries via a data-driven Adaptive Phase Partition module, identifies diagnostically critical phases with attention, and extracts disease-related connectivity within each phase using an Incremental Graph Structure Generator regularized by binarization, temporal smoothness, and sparsity; a spatio-temporal supervised contrastive learning approach then leverages diagnosis-relevant spatio-temporal patterns to refine the similarity metric between samples and build a well-structured semantic space. Experiments on ASD, BD, and MDD validate its effectiveness, and the discovered critical phases and subnetworks provide interpretable evidence consistent with prior neuroimaging findings.
Fig. 1. Overview of the proposed BrainSTR framework. Given a subject’s BOLD signal X ∈RT ×N, where T and N denote the numbers of time points and ROIs, respectively, Adaptive Phase Partition (APP) infers state-consistent phase boundaries to construct phase-wise FCs {At}W
· Page 3Interpretation
BrainSTR decomposes dynamic brain network modeling into phase partitioning, critical-phase identification, and within-phase connectivity extraction, yielding spatio-temporal interpretability. Prior dynamic functional connectivity work often focuses on overall diagnostic performance, whereas this framework explicitly targets when discriminative disease signatures emerge and where they reside in the connectivity topology. The abstract describes the combination of an Adaptive Phase Partition module, attention, and an Incremental Graph Structure Generator, and reports validation on ASD, BD, and MDD; specific metrics and sample sizes are not given in the abstract.
The Incremental Graph Structure Generator extracts disease-related connectivity within each phase under three regularizations: binarization, temporal smoothness, and sparsity. This design responds to diagnostic signals being subtle and sparsely distributed across time and topology amid pervasive non-diagnostic connectivities, constraining extracted structure to be sparse and temporally coherent. The abstract explicitly lists the three regularization terms and their targets, which is a method-level design statement; the magnitude of gains over baselines is not quantified in the abstract.
Spatio-temporal supervised contrastive learning uses diagnosis-relevant spatio-temporal patterns to refine the similarity metric between samples, capturing more discriminative features and constructing a well-structured semantic space. Unlike conventional contrastive learning, diagnosis-relevant spatio-temporal patterns are injected as supervision into the similarity metric, so the representation space serves both discrimination and interpretability. The abstract states the module's motivation and goal and reports validation on ASD, BD, and MDD; specific numbers from ablations and comparisons are not presented in the abstract.
The discovered critical phases and subnetworks are consistent with prior neuroimaging findings, providing interpretability support for model outputs. This aligns the model's selected temporal phases and connectivity subnetworks with existing neuroimaging literature rather than reporting classification accuracy alone. The abstract phrases the agreement as 'consistent with prior neuroimaging findings', a qualitative comparison; the specific brain regions, phases, or references are not listed in the abstract.
Perspective
The work targets dynamic brain network analysis where spatio-temporal interpretability is needed, suited to researchers and clinical imaging directions that already have dynamic functional connectivity data and want to localize discriminative phases and subnetworks. It enables follow-up work to test whether critical phases are stable and subnetworks reproducible, and to apply the phase-partitioning and sparse graph extraction ideas to other dynamic graph or temporal signal modeling tasks.
The abstract does not give specific evaluation metrics, sample sizes, baseline comparisons, or ablation results, so the relative contribution of each module cannot be judged from the abstract; the concrete correspondence behind 'consistent with prior neuroimaging findings' is also not listed. In addition, the available text is the abstract plus page navigation information, and figures and experimental details from the body are not included, so assessing empirical strength still requires the original paper.
