BrainHO replaces fixed brain atlases with learnable subgraphs, reaching 69.68% accuracy on ABIDE from PCC input alone while localizing cross-network disease subgraphs
Synopsis
The work proposes Brain Hierarchical Organization Learning (BrainHO), which uses learnable subgraph and graph tokens to aggregate brain regions bottom-up via hierarchical attention driven by node feature affinity, combined with a subgraph orthogonality constraint and a hierarchical consistency constraint; on ABIDE (N=1009, 516 ASD/493 healthy controls) and REST-meta-MDD (N=2380, 1276 MDD/1104 healthy controls) it attains the highest accuracy (69.68% and 64.71%) and sensitivity (73.11% and 67.43%) using only the static PCC connectivity matrix, while visualizing disease-related subgraphs that partly overlap predefined networks such as the SMN and DAN and partly span multiple predefined networks.
Interpretation
BrainHO is proposed to dynamically aggregate brain regions through learnable subgraph tokens driven by node feature affinity, replacing the rigid partitions of predefined functional atlases such as Yeo7 and Smith10, and thereby capturing cross-network interactions that fixed atlases split apart. Prior methods such as Com-TF, CAGT, and DHGFormer rely on predefined sub-networks and model intra- and inter-subnetwork relations separately, implicitly assuming strict functional boundaries; this work instead learns hierarchical organization from intrinsic features without predefined sub-network labels. The paper shows average Pearson correlations between brain regions on ABIDE that reveal strong interactions across predefined sub-networks, and gives the concrete example that Frontal_Inf_Oper_L and Temporal_Sup_L are hard-coded into separate communities (VAN and SMN) in Yeo7, so the fixed atlas cannot capture that interaction.
A hierarchical attention mechanism is designed to aggregate information bottom-up across node, subgraph, and graph levels, using Sparsemax in the node-to-subgraph stage to produce sparse attention that filters noise and yields interpretable sub-networks, and Softmax in the subgraph-to-graph stage to densely aggregate subgraph information. Unlike flat attention in a standard Transformer, this mechanism explicitly introduces learnable subgraph tokens and a global graph token as intermediate levels; in ablation, the Baseline and the w/o HA variant reach only 61.55% and 61.74% accuracy, below the full model. The method presents equations for node-to-subgraph and subgraph-to-graph attention, and supports the role of hierarchical attention with the Baseline and w/o HA ablation comparisons on ABIDE.
A subgraph orthogonality constraint and a hierarchical consistency constraint are introduced: the former applies a cross-entropy loss over the cosine similarity matrix of L2-normalized subgraph tokens to encourage diverse, non-redundant sub-networks, while the latter lets the aggregated graph token act as a teacher and a node-level auxiliary classification head as a student, aligning global semantics with local node representations via KL divergence. These two constraints target possible redundancy and feature collapse among learnable subgraphs, and the lack of high-level semantic guidance for local node features, adding regularization designs beyond existing predefined-atlas methods. Ablation shows accuracy drops to 65.61% without the orthogonality constraint, to 67.20% without the auxiliary supervision Laux, and to 68.89% without the consistency constraint Lhc, all below the full model, providing controlled evidence for both constraints.
The method achieves the highest accuracy and sensitivity on ABIDE and REST-meta-MDD, and visualizes disease-related sub-networks that are both functionally consistent and cross-network. BrainHO uses only the static PCC connectivity matrix, whereas ALTER, DHGFormer, and LHDFormer additionally incorporate raw BOLD temporal input; on REST-meta-MDD, LHDFormer has a marginally higher AUC, but BrainHO has the highest accuracy and sensitivity. Results are based on 5-fold stratified cross-validation with reported means and standard deviations; for interpretability, the No. 7 sub-network for MDD and the No. 5 sub-network for ASD overlap the SMN and DAN by 66.7% and 68.4% of their nodes respectively, the ASD No. 6 sub-network includes nodes such as Fusiform_L and Amygdala_R, and the MDD No. 1 sub-network includes the Hippocampus and Cingulum_Ant.
Perspective
The framework targets brain disorder classification settings that take an fMRI-derived static PCC connectivity matrix as input, evaluated on the two public datasets ABIDE and REST-meta-MDD with 5-fold stratified cross-validation, and is suited to researchers who want ASD and MDD discrimination and sub-network localization without introducing raw BOLD temporal input. Its interpretability output is presented by mapping attention weights onto predefined functional networks and AAL116 names, which supports follow-up mechanistic work on cross-network anomalies.
The paper reports cross-validation results from a single study on two datasets and does not yet constitute clinical consensus; the link between cross-network sub-networks and disease labels comes from visual interpretation of attention weights, and its stability awaits testing on more independent data. In addition, the text describes Figures 1 and 3 in prose, so the specific sub-network mapping matrices and node-level details require the original figures, and the visualization conclusions are hard to fully reproduce from the text alone.
