Removing the overlapping-window timing shortcut lets non-invasive brain-to-text reach 36.6% word error rate with five observations per word
Synopsis
The work shows that most of the reported gain from jointly decoding all words in a sentence (d'Ascoli et al., 2025) is reproducible on synthetic signals containing no brain information (22.0% versus 22.3% on real MEG), because neighbouring three-second word-aligned windows overlap and implicitly leak word duration; decoding words independently instead (SimpleB2T) makes aggregation across observations and an LLM prior effective, reaching 36.6% word error rate with five observations per word on a clinically motivated benchmark.
Interpretation
Most of the reported gain from jointly decoding words does not require brain activity and can be reproduced from word-duration information exposed by overlapping windows. That gain had been attributed to jointly modelling neural responses within a sentence; the authors use a synthetic continuous signal that preserves the overlap structure but carries no stimulus information to localise the gain to the input construction itself. On subject 0 of LibriBrain100, joint decoding of real MEG reaches 22.3% balanced word accuracy versus 9.5% for isolated decoding; a synthetic signal with the same overlap reaches 22.0%, dropping to 5.8% when overlap is removed, while a timing-only control reaches 22.9%. The effect replicates across the other 32 subjects and on the Armeni et al. (2022) and Le Petit Prince datasets.
The leakage arises because neighbouring windows share samples, which directly reveals the onset-to-onset interval, and that interval is strongly related to word duration. The authors frame this explicitly as an instance of shortcut learning and give a solvable form: shared samples in the overlapping region satisfy the same sampling relation, so minimising a sum of squared differences recovers the interval. In the natural speech data, 99.9996% of adjacent word pairs have overlapping windows sharing 90.6% of their samples on average, and across 122,120 adjacent word pairs the onset-to-onset interval correlates strongly with the annotated duration of the preceding word; joint decoders align more strongly with a duration-based word posterior than isolated decoders.
Decoding words independently makes two established strategies, aggregating repeated neural responses and adding an LLM linguistic prior, substantially more effective. Under joint decoding, aggregation gives little benefit and LLM rescoring performs close to or worse than an LM-only baseline; once the shortcut is removed, neural evidence and the linguistic prior become complementary. On the 200-sentence clinical communication benchmark with a 92-word vocabulary, LLM rescoring at five observations reduces WER from 74.3% to 36.6%, and going from one to five observations reduces WER from 65.6% to 36.6%; the neural decoder alone gives 74.3% and the LM alone 73.8%, while noise inputs matched to MEG statistics give 97.1% and shuffled predicted embeddings give 88.3%.
The shortcut extends beyond the original work, though it is not universal across aligned-window tasks. The authors broaden the analysis to several later non-invasive brain-to-text methods and benchmarks and run the same synthetic control for Brain2Qwerty. Of the nine datasets analysed by d'Ascoli et al. (2025), only LittlePrinceRead and Nieuwland are unaffected because their reading protocols give each word the same amount of time, and these are also the only two datasets where joint decoding did not improve over independent decoding; for Brain2Qwerty, synthetic inputs perform only slightly above chance (6.5% validation and 6.8% test balanced character accuracy versus 37.0% and 34.9% with real MEG), so timing cannot explain its performance.
Perspective
The result applies to word-aligned non-invasive brain-to-text: word onset times are known at decoding time, and evaluation is confined to a constrained clinical communication domain with a 92-word vocabulary and a 200-sentence benchmark. The authors note the benchmark is built from the public LibriBrain100 data, is not clinically validated, and should be evaluated with users before any real-world use. Methodologically, SimpleB2T keeps the word-level encoder, semantic targets, and contrastive training framework of d'Ascoli et al. and only replaces sentence-level joint encoding with independent per-word encoding, reducing the model from roughly 200 million to about 20 million parameters; aggregating observations requires no retraining of the neural decoder, so collecting more observations becomes a way to trade recording burden for reliability. The authors also note that when word onsets are unknown this kind of leakage is not an issue, but that is a substantially different and more difficult task they do not evaluate, and extending the results to internally generated speech such as imagined speech is a key step towards practical non-invasive communication interfaces.
Several open questions remain for a careful reader. First, the evaluation concerns perceived speech rather than imagined speech, and the authors explicitly do not present the method as a practical BCI. Second, the word-alignment assumption means word onsets are known, which differs from jointly solving speech segmentation and word decoding. Third, the clinical communication benchmark is not clinically validated, and its sentences were constructed with generative-AI assistance and reviewed by the authors, so performance in real patient settings remains to be evaluated with users. Fourth, the authors identify a remaining bottleneck in selecting among beam-search candidates, since better sentences are often already present in the beam, and reducing the number of required observations is likewise unresolved. Fifth, although this is a full-text reading, several results tables in the supplied text do not include their numeric values, so some ablations and comparisons can only be summarised from the prose rather than from table numbers.
