Skip to main content
Back to timeline
arXivSource publication:

Treating EEG masking geometry as the only variable: 58 pre-trained models point to a moderate spatial radius with short temporal blocks, and expose a JEPA-specific bias-inflation collapse

Synopsis

The study unifies EEG self-supervised masking strategies into a three-parameter framework of spatial radius, temporal length and mask ratio, trains 58 models under a fixed REVE-Small backbone and corpus across MAE and JEPA, evaluates them with a linear probe on the 12 OpenEEGBench datasets, and finds that both frameworks agree on a moderate spatial radius with short temporal blocks as the optimum while identifying a JEPA-specific bias-inflation collapse at full-channel masking.

AI-generated editorial illustration: What masking geometry works best for EEG foundation models?

Interpretation

The paper introduces a three-parameter masking framework (spatial radius, temporal length, target mask ratio) that recovers random patch masking, whole-channel masking, contiguous time-window masking and spatio-temporal block masking as boundary cases and continuous interpolations of one family. Previously each new EEG foundation model bundled a new masking strategy with a new backbone, corpus and objective, so masking was never ablated in isolation; this parameterization makes those shapes comparable under one formalism. A formal definition plus a mask sampling algorithm, with appendix simulations confirming the realised mask ratio matches the target and that empirical token-distance distributions track the analytical reference.

MAE and JEPA agree on the optimum, a moderate spatial radius with short temporal blocks, and on shared failure modes: single-channel blocks, full-channel blocks, and most long temporal blocks. This is the first replication of the same masking conclusion across two self-supervised objectives under a single fixed pipeline, indicating the optimum is not an artefact of one objective. 58 pre-trained models, 12 OpenEEGBench datasets, linear probes over 5 random-projection seeds, and two-level bootstrap statistical testing.

Outside the failure modes, downstream performance is largely insensitive to the precise mask: 11 MAE and 9 JEPA configurations are statistically indistinguishable from their framework's best, and the within-framework spread exceeds the cross-framework gap at the optimum. It separates 'masking is the dominant lever' from 'within the working regime the choice is free', giving practitioners an actionable default rather than a single optimum. Hierarchical bootstrap over within-dataset ranks and pairwise dominance probabilities, with reported best-minus-worst normalised score spreads per framework.

It identifies a JEPA-specific 'bias-inflation collapse': when every channel of the same time-patch is masked jointly, the encoder converges to a position-only lookup table, the per-channel feature mean inflates while the input-driven residual shrinks, and standard collapse detectors (per-dimension variance, loss divergence, trivial-constant collapse) never fire. This failure mode is not documented in prior literature and differs from low-rank collapse, since the effective rank actually grows, so rank-based probes are fooled as well. Bias and residual magnitudes plus intra-channel cosine similarity are tracked across pre-training epochs on physionet and bcic2a and replicated on five subjects, with a mechanistic appendix and ruled-out hypotheses.

Perspective

The result targets practitioners pre-training EEG foundation models with a masked-prediction objective, applies to pre-training corpora dominated by clinical EEG and BCI paradigms, and the recommended geometry has been validated on backbones beyond REVE-Small, namely CBraMod and LaBraM. It supplies a working order: fix the mask before scaling compute, since at a fixed budget a poor mask costs up to a normalised score. For teams using a JEPA-style objective, the directly usable rule is never to mask every channel of the same time-patch jointly, and to prefer a moderate spatial radius.

The core grid uses a single backbone and a single pre-training corpus, so transfer to other corpora, hyperparameters and consumer-grade EEG remains to be established; the mask ratio is fixed in the core grid, ratio and geometry interact, and a full joint sweep remains to be done. One of the 12 downstream datasets, PhysioNet-MI, is part of the pre-training corpus while the other eleven are unseen, and OpenEEGBench overlaps with the pre-training corpora of other recent EEG foundation models, so collective benchmark overfitting cannot be ruled out. Within the mechanistic account of bias-inflation collapse, whether VICReg-style regularisers would actually prevent the solution, and whether distributional regularisers such as SIGReg would rule out the lookup-table fixed point, are left untested, and a reliable in-training detector has not been built. In addition, several tables in the loaded text render their numeric cells as placeholders, so specific scores and compute figures cannot be checked item by item here.

Sources