AlloPool: a graph neural network framework that infers protein allostery from molecular dynamics simulations
Synopsis
The study presents AlloPool, a graph neural network framework that combines temporal attention with iterative edge pooling to learn minimal, time-evolving residue interaction networks from equilibrium and non-equilibrium molecular dynamics trajectories, reconstructing trajectories at sub-angstrom RMSD across Pin1, the GAIN mechanosensor domain, dopamine D2 and beta-1 adrenergic receptors, the SdrG adhesin, PDZ3, and the Engrailed homeodomain, and using those networks to map allosteric pathways, predict ligand pharmacology, mechanical loading states, and mutation effects.
Interpretation
AlloPool learns minimal, time-evolving interaction graphs from trajectories through temporal attention and iterative edge pooling, enabling reconstruction of protein dynamic trajectories. Unlike NRI, which relies on static, fully connected graphs, this framework explicitly models time-dependent conformational change and prunes non-contributing edges. Benchmarked against NRI on Pin1, GAIN, D2, and B1AR, AlloPool achieved sub-angstrom reconstruction RMSD while NRI RMSD errors ranged from 1.4 to 5.72 Å; ablation on GAIN raised RMSD from 0.89 Å to 1.3 Å without temporal attention and to 2.1 Å without pair attention.
The learned directed, asymmetric edge networks reconstruct force propagation pathways under mechanical load and distinguish the unloaded GAIN state from two mechanically loaded states. Correlation-based methods and PRS, which rests on linear response theory, struggle with induced fit or mechanical unfolding, whereas AlloPool's pooled-edge frequencies shift systematically with applied force. GAIN steered MD at 0.1 and 1 nm/ns pulling speeds produced two loaded states; reconstructed force pathways aligned with correlation-based results, and SdrG-fibrinogen peptide rupture force distributions and force-extension curves agreed with prior reports.
Pooled-edge features unsupervisedly resolve GPCR functional states and predict ligand efficacy. Compared with contact persistence metrics commonly used in GPCR structural analysis, pooled-edge features correlate more strongly with ligand efficacy. Clustering a cumulative 2-microsecond D2 receptor trajectory yielded three dynamic modes corresponding to inactive, active, and intermediate states; in B1AR, persistence metrics correlated weakly with efficacy while pooled-edge features produced many edges with R2 > 0.9, with predictive residues on TM3, TM5, and TM6.
The learned interaction networks predict the impact of allosteric mutations on binding function and reveal dynamic consequences of missense mutations that structure prediction misses. Where AlphaFold2 and AlphaFold3 produced nearly identical wild-type and L16A Engrailed homeodomain structures (RMSD 0.339 Å), AlloPool detected a marked weakening of interaction density between H1 and the helical bundle from dynamics. In PDZ3, residues in the 30 highest-coupled edges overlapped strongly with known gain- and loss-of-function mutation positions; across eight PDZ variants, loss-of-function variants showed fewer paths or weaker coupling and gain-of-function variants the opposite, except the strongly destabilized N363F; the Engrailed mutant was sampled for approximately 6 microseconds of equilibrium dynamics.
Perspective
The framework targets systems with existing molecular dynamics or steered molecular dynamics trajectories, applies to equilibrium and non-equilibrium simulations, and covers small soluble domains, membrane receptors, mechanosensors, and multi-domain proteins; its output is a system-specific minimal interaction network and dynamic modes that can be used to locate state-specific allosteric sites, predict ligand efficacy and mutation effects, and generate candidate hypotheses for experimental validation and protein engineering.
The model currently does not encode amino acid identity or chemistry and relies only on spatial information from trajectories, so interaction networks are system-specific and cross-protein generalization requires retraining; simulations of the Engrailed mutant were not long enough to reproduce the experimentally observed large-scale rearrangement, so the model detected weakened interaction networks rather than the conformational displacement itself; additionally, this is a full-text parse in which some figures appear as supporting material, so verifying specific values and statistical details still requires consulting the original figures and data files.
