KD-Brain guides subnetwork interaction modeling with semantic priors and a pathology-consistent constraint, beating 12 baselines on ASD, BD, and MDD diagnosis while yielding interpretable functional pathways
Synopsis
The work proposes KD-Brain, a prior-informed graph learning framework that injects disorder-specific semantic priors into the attention Query (Semantic-Conditioned Interaction) and aligns learned subnetwork interaction distributions with clinical priors via a KL-divergence Pathology-Consistent Constraint, achieving better performance than 12 baselines on the ABIDE (NYU) ASD task and single-center BD and MDD tasks while producing functional pathways and critical brain regions consistent with psychiatric pathophysiology.
Fig. 1. Illustration of the KD-Brain.
· Page 3Interpretation
KD-Brain formulates the brain network as a heterogeneous graph, extracts subnetwork-level topological embeddings with a bidirectional convolutional spatial encoder, learns inter-subnetwork interactions via Transformer attention, and reads out interaction-strength distributions as an interpretable output. Compared with GNN methods that fit local topology within fixed structures (e.g., HeBrainGNN, MVS-GCN) and Transformer methods with unguided attention (e.g., BNT, Com-BrainTF, CAGT), this work explicitly models higher-order interactions between functional subnetworks. Evaluated on ABIDE (NYU, 74 ASD patients and 98 healthy controls) and a single-center dataset (246 healthy controls, 151 MDD patients, 126 BD patients), compared against 12 baselines, reporting mean±std for ACC and AUC.
Semantic-Conditioned Interaction encodes disorder-specific subnetwork descriptions (e.g., DMN 'hypoconnectivity' in ASD versus DMN 'hyperconnectivity' in MDD) with BioMedBERT and injects them into the attention Query, acting as a functional-identity positional encoding. Unlike prior work that uses only classification labels or one-hot/3D-coordinate positional encodings, this method distinguishes the same subnetwork across disorders using its neurocognitive role and disease-related pathological alteration. Ablation shows replacing semantic interaction learning with a standard GAT causes a marked drop (e.g., 5.0%↓ on ASD); removing either the semantic prior Hsp or the pathological constraint Lpmc degrades performance; Fig. 2 reports significant Pearson correlations among disorder-agnostic prior embeddings (P < 0.05), and GPT-4 and DeepSeek-R1 yield a consistent ranking of interaction strengths.
The Pathology-Consistent Constraint uses KL divergence to pull the learned subnetwork interaction distribution Psni toward an LLM-prompted clinical prior distribution Pprior, with the total loss being a weighted sum of cross-entropy and the knowledge regularization term. Treating clinical priors as a training-time regularizer rather than only post-hoc interpretation keeps interaction patterns neurobiologically plausible under limited samples. Ablation shows removing Lpmc degrades performance across all three tasks; the text reports that GPT-4 and DeepSeek-R1 produce a consistent prior ranking.
The model's multi-level biomarkers align with psychiatric literature: 'DMN→CEN→DMN' for ASD, 'CEN→CEN→CEN' for BD, and SN-initiated pathways (e.g., SN→CEN→CEN) for MDD; the Thalamus and Orbitofrontal Cortex rank as top biomarkers for both BD and MDD. These pathways and regions are derived directly from attention coefficients and cross-order joint probabilities, offering computational evidence for shared neuropathology. Pathway scores are defined as the joint probability of sequential interactions across orders (Score = ∏ α), and region rankings come from model visualization, which the authors note align with established neurobiological literature.
Perspective
The results target psychiatric classification settings that take fMRI functional connectivity as input, parcellate with the AAL atlas, and group regions into three subnetworks (DMN, SN, CEN), covering ASD, BD, and MDD diagnosis tasks; they are informative for researchers and clinical research directions seeking to inject domain priors into small-sample medical graph data while obtaining interpretable interaction pathways.
Readers should still watch: the priors are LLM-prompt generated, and their content and ranking stability across prompts or models is supported in the text only by the consistent ranking of GPT-4 and DeepSeek-R1; ABIDE uses only the NYU single site and BD/MDD come from a single center, so cross-site and cross-cohort generalization remains to be tested; pathway and region rankings come from visualization of attention coefficients and joint probabilities, whose statistical robustness needs validation on more independent data; this is a fast parse, and the specific values in Fig. 1, Fig. 2, and Fig. 3 as well as hyperparameter settings (such as the choices of λsp, β, and q) are not fully given in the text, and these details affect reproduction.
