Personalized EEG models of epilepsy and non-epilepsy show similar predictive fit but markedly different learned latent dependency structure
Synopsis
Using frozen representations from a lightweight CNN–Transformer EEG foundation model pretrained on TUEG, the study maps TUEP segments into a shared latent-state space and fits sparse multinomial logistic transition distributions independently per subject to obtain personalized latent transition-dependency graphs; it finds that epilepsy and non-epilepsy subjects are comparably predictable within-subject yet epilepsy models retain substantially denser learned dependency structure, with graph-derived features retaining moderate group discrimination.
Figure 1 : Pipeline for personalized latent-dynamics modeling from scalp EEG. Non-overlapping 30-s EEG segments are encoded using our frozen EEG foundation model and mapped to shared latent-state spaces using global k k -means at k = 4 k{=}4 and k = 6 k{=}6 . Each segment is assigned a discrete state label, producing a subject-specific state sequence. The sequence is modeled independently for each subject using mLTD with group-lasso regularization across temporal lags, yielding a sparse, directed, weighted transition-dependency graph W n ∈ ℝ k × k W_{n}\in\mathbb{R}^{k\times k} .
arXivInterpretation
The work introduces and applies a pipeline that separates predictive validity from characterization of learned dynamical structure: frozen EEG foundation-model representations, a shared latent-state space, and per-subject mLTD fits yielding sparse directed weighted latent transition-dependency graphs. Prior clinical AI evaluation focuses heavily on predictive performance; this work treats “how well the model predicts” and “how the model organizes temporal dependencies” as two complementary evaluation axes, explicitly distinguishing state occupancy from state-to-state transitions. The pipeline is specified in detail: the encoder was pretrained by self-supervised reconstruction on 2,259,998 non-overlapping 30-s TUEG segments (about 18,800 hours) and frozen for TUEP; each 30-s segment maps to a 128-dimensional representation; a single global k-means yields the shared state space; maximum lag and sparsity are selected per subject via 5-fold time-series cross-validation.
Despite nearly identical predictive fit, epilepsy subjects retained substantially denser learned dependency structure: at the primary resolution, mean nonzero dependencies were 10.01 vs. 5.67 with normalized densities about 62.6% vs. 35.4%; at the finer resolution, 19.90 vs. 13.46 with normalized densities about 55.3% vs. 37.4%. This provides an empirical fit–structure dissociation: the two groups are comparably predictable under their personalized models while organizing temporal dependencies differently, showing that predictive equivalence is insufficient to infer equivalent internal dynamics. Group analyses use 99 epilepsy and 99 non-epilepsy labeled subjects with Mann–Whitney tests; the same group-level pattern appears at two latent-state resolutions, and appendix lag robustness checks show predictive fit remains matched across lags.
Features derived from the dependency graphs retained moderate epilepsy versus non-epilepsy discrimination, with scalar summaries, graph-derived features, and vectorized full-matrix features yielding moderate AUROCs under stratified 5-fold subject-wise cross-validation. This suggests clinically associated information may reside in how learned latent dynamics are organized, not only in overall predictive fit; the authors stress this is not diagnostic performance. Classification uses L2-regularized, class-balanced logistic regression with stratified 5-fold subject-wise cross-validation; sensitivity analyses excluding null (all-zero) graphs kept scalar group AUROC above chance, indicating densification and discrimination are not solely an artifact of all-zero graphs.
The results are positioned as an intermediate evaluation step toward patient world models: structural characterization should precede extending personalized dynamical models toward intervention-aware reasoning. The authors outline conditioning dynamics on a measured intervention or perturbation and characterizing intervention-associated structural change, while noting this requires longitudinal or intervention-conditioned data and appropriate causal assumptions. This positioning rests on the present observational case–control design; the authors state the learned dependencies should be interpreted as research descriptors of model-learned latent dynamics rather than diagnostic biomarkers, causal brain graphs, or physiological connectivity maps.
Perspective
The work is aimed at research settings: using scalp EEG with epilepsy as a clinical contrast, it evaluates whether personalized latent dynamical models carry clinically associated information beyond predictive fit. It is suited to researchers and engineering teams who need to check whether clinically associated dynamical structure is present in a learned latent space before incorporating such representations into downstream clinical AI or intervention-aware models. The authors state the learned dependency graphs should be treated as research descriptors of model-learned latent dynamics rather than diagnostic biomarkers, causal brain graphs, or physiological connectivity maps, and should not be used for autonomous diagnosis, treatment selection, or intervention planning without prospective clinical validation.
The learned k-means states are unsupervised and lack established clinical or physiological semantics; the analysis is observational and case–control, so differences in dependencies represent predictive temporal dependencies rather than causal disease mechanisms or treatment effects. Higher null-model prevalence among non-epilepsy subjects at the primary resolution may partly contribute to the observed density difference, although non-null sensitivity analyses indicate clinically associated structural information is not solely attributable to all-zero graphs. The personalized analysis uses representations from a single frozen EEG foundation model, and cross-encoder concordance of subject-specific dependencies remains to be established. Prospective, test–retest, and longitudinal analyses will be required to establish the stability and clinical generalizability of the learned dynamical structure.
