In motor imagery BCIs, seven out-of-distribution detection methods failed due to intrinsic EEG variability, but Deep Ensembles and MC-Dropout reached up to 0.7 OOD detection ability for subjects with high on-task performance
Synopsis
This study used a Leave-One-Class-Out out-of-distribution detection setup in motor imagery BCIs, training a model on some classes and observing whether an unfamiliar movement class can be detected via increased uncertainty; it found that because of the high intrinsic variability of EEG signals, many users show higher uncertainty for familiar in-distribution classes than for out-of-distribution classes, so many OOD detection methods that perform well in other machine learning domains prove ineffective here, yet OOD detection performance correlates with on-task performance, and Deep Ensemble and MC-Dropout models achieved on-task AUROC above 0.9 and OOD detection ability up to about 0.7, showing that rejecting unfamiliar cognitive states becomes feasible when task performance is high.
Interpretation
The work proposes and demonstrates a Leave-One-Class-Out OOD detection experimental paradigm for studying whether models are robust against unfamiliar motor imagery actions not seen in training data. The authors state this is the first benchmark in motor imagery BCI that evaluates this type of OOD data. Based on an experimental setup that trains on some classes and then observes an unfamiliar movement class; the abstract describes the paradigm design but does not provide subject counts or dataset details.
In motor imagery BCIs, many OOD detection methods that have shown good performance in other machine learning domains prove ineffective at identifying unfamiliar motor imagery patterns. The cause is the high intrinsic variability inherent in EEG signals, which can make uncertainty for familiar classes higher than for out-of-distribution classes for many BCI users. The study tested seven different OOD detection methods and one more method claimed to boost OOD detection quality; the abstract reports that these methods performed poorly in this setting.
OOD detection performance is correlated with on-task performance, and Deep Ensemble and MC-Dropout models achieved on-task AUROC above 0.9 and OOD detection ability up to about 0.7 for models and subjects with high task performance. This shows that rejecting unfamiliar cognitive states becomes feasible when task performance is high, offering a path to improve the overall safety and reliability of BCIs. The abstract reports the above AUROC values and the correlation, but does not give subject numbers, statistical tests, or confidence intervals.
Perspective
This work is aimed at researchers and system designers of motor imagery BCIs, and applies to settings where a model is trained on some motor imagery classes and then faces unfamiliar movement classes not seen in training. The Leave-One-Class-Out OOD detection paradigm it proposes can serve as an experimental framework for subsequent studies of model robustness, and it suggests that rejecting unfamiliar cognitive states is feasible when task performance is high, pointing toward improved BCI safety and reliability.
The abstract does not state the number of subjects, dataset sources, or statistical tests, so the robustness of the reported OOD detection ability up to about 0.7 still needs to be assessed with the full text. In addition, whether the correlation between OOD detection performance and on-task performance is stable across different subject groups, and whether methods beyond Deep Ensembles and MC-Dropout can improve results when task performance is lower, are questions worth watching. Because only the abstract was read here, figures and detailed results in the full text could not be included in this summary.
