ZeroMAG generates multimodal adapters zero-shot from unlabeled calibration signals, raising balanced accuracy of frozen EEG foundation models by 7.22 points on average across six held-out datasets
Related research and updatesSynopsis
ZeroMAG introduces a zero-shot multimodal adapter generation framework that, with a frozen EEG encoder and prediction head, uses only unlabeled target calibration recordings and task context to generate adapter weights through a configuration-invariant 1+N adapter, a modality–subject–task condition, and a function-constrained latent space learned from source adapters; across six held-out target datasets and three EEG foundation model backbones it improves balanced accuracy by 7.22 percentage points over EEG-only inference and 4.89 points over direct weight regression, coming within 0.50 points of supervised multimodal adaptation on average, with no target labels or target-side optimization.
Figure 1: Multimodal EEG landscape and representative results. (a) Physiological information captured by common companion signals. (b) Publication trends, task-wise multimodal prevalence, and companion modality availability, based on prior surveys ( Lee et al., 2025 ; Yeung & Chu, 2022 ) . (c) Balanced accuracy on six held-out target datasets, comparing three EFM backbones with and without ZeroMAG against SleepFM and PhysioOmni.
arXivInterpretation
ZeroMAG frames multimodal extension as personalized parameter generation: given unlabeled target recordings and task context, it directly generates a target-specific adapter while keeping the EEG foundation model frozen. Unlike parameter-efficient fine-tuning and test-time adaptation that require target-side optimization, and unlike multimodal physiological foundation models that require multimodal pretraining, adaptation is amortized into an offline-trained generator, and deployment needs a single forward generation pass. Across six held-out target datasets and three backbones (CBraMod, CSBrain, CodeBrain), covering 18 backbone–dataset settings, ZeroMAG outperforms EEG-only inference, mean source weights, nearest source weights, and Direct MLP in every setting, by 7.22, 6.78, 6.00, and 4.89 percentage points on average, respectively.
A configuration-invariant 1+N adapter treats the frozen EEG representation as a fixed anchor and attaches one homogeneous branch per available companion modality (temporal–spectral encoding, feature projection, shared coordination), while missing modalities create neither parameter blocks nor placeholder latent tokens. A single generator covers heterogeneous modality configurations without per-dataset redesign and without fixing the supported modality set in advance. Ablations show a shared modality encoder lowers accuracy by 4.16 points, a temporal-only branch by 8.12 points, and replacing learned coordination with mean fusion by 6.58 points, supporting modality-specific encoding, temporal–spectral processing, and learned fusion.
A modality–subject–task condition infers target-specific information from unlabeled recordings without dataset identifiers; the subject factor is inferred from the observed recording rather than retrieved from a subject identifier. Modality information alone cannot distinguish subject-specific physiology under the same modality combination, nor the same physiological information serving different tasks. Removing the subject factor and retaining only task and modality lowers balanced accuracy by 11.09 points; shuffling subject conditions lowers it by 9.24 points even though the representations retain the same marginal distribution and dimensionality.
Function-constrained two-stage generation (function-preserving representation learning plus function-constrained conditional generation) operates in a block-aligned latent space and supervises predictive agreement through the frozen EFM rather than minimizing parameter error alone. Close parameter reconstruction does not guarantee close predictions; the design extends the generation target from weight space to predictive behavior. Removing Stage 1 functional supervision reduces reconstruction error by 6% yet increases generated-weight error by 5% and prediction KL by 31%, with B-ACC falling 1.15 points; removing Stage 2 functional supervision raises prediction KL by 48% and lowers B-ACC by 1.58 points; removing both lowers B-ACC by 2.87 points and raises prediction KL by 82%.
Perspective
The result targets settings with a supplied fine-tuned EEG foundation model and a compatible prediction head: the target side needs only unlabeled calibration signals and a known task to generate a plug-and-play adapter, with no target labels, pseudo-labels, gradients, or iterative optimization. Evaluation spans sleep staging, emotion recognition, cognitive load, motor imagery, and freezing-of-gait detection, with companion signals including EOG, EMG, ECG, eye movement, acceleration, and skin conductance; on unseen combinations of known modalities and unseen modality types it recovers 89.7% and 82.6% of the supervised multimodal gain, respectively. Performance saturates at approximately 64 unlabeled calibration windows, and using 50% of the source adapter bank retains 97.2% of the full-bank multimodal gain; target-side condition construction and adapter generation take about 2 minutes versus about 15 minutes for supervised target adaptation. The method depends on the coverage and quality of the source adapter bank, and offline training cost is amortized across targets.
The ablation variants are separately trained systems, so their decreases are incremental differences relative to the full model rather than an additive causal decomposition; the emotion-recognition target SEED-V shows both smaller absolute improvement and lower supervised-gain recovery than other targets, and the text does not identify the source of this task-dependent gap. In the cross-paradigm comparisons, PhysioOmni, SleepFM, and Brant-X use their own native backbones and different multimodal learning procedures, so backbone architecture, pretraining exposure, and target adaptation are not isolated as single controlled factors. Reported standard deviations summarize seed-to-seed variation, not confidence intervals for pairwise differences, and overlapping whiskers establish neither significance nor equivalence; the 0.50-point supervised gap is a small residual in displayed averages, not a formal equivalence result. Functional matching transfers the source reference adapter's behavior and its errors, so it is not an unsupervised proof of correctness. Calibration window counts describe the number of calibration units rather than a universal duration, since window duration and sampling rates differ across tasks. A joint optimum over source-bank size and calibration budget is not provided. In addition, the main result tables are empty in the loaded text, so per-dataset numbers can only be understood through the appendix aggregates and the averages stated in the main text.
