‘Multifunctional’ brain implant translates speech and gestures in real time
Synopsis
A Nature news report describes a proof-of-concept study published in Nature Neuroscience in which a single surgically implanted 253-electrode array covering a fairly large area of the sensorimotor cortex, combined with artificial intelligence, simultaneously decoded attempted phrases and attempted or imagined gestures in two participants with impaired speech and movement after a brainstem stroke or with amyotrophic lateral sclerosis, producing on-screen text and driving a personalized animated avatar within seconds of the user’s intent, thereby translating both verbal and non-verbal communication through one implant.
Brain–computer interfaces decode users’ brain signals. Credit: Kevin Frayer/Getty
PubMedInterpretation
A single implant can read brain signals related to both speech and body movement and output text and avatar motion at the same time. Most prior BCI studies restored one function at a time, such as talking, cursor control or a robotic arm; this work handles verbal and non-verbal communication in parallel within one system. Proof of concept with two participants, using a surgically implanted 253-electrode array on the cortical surface covering a fairly large sensorimotor area.
The study offers a workable path to decoding speech and gestures together, even though the neural signals involved overlap and activity patterns during simultaneous speech and gesturing differ slightly from speech-only or gesture-only tasks. It treats signal overlap as a design target, using broad cortical coverage to capture signals related to both speech and body movements. Evidence comes from neural activity recorded while the two participants attempted speech, attempted or imagined gestures, and combined the two, such as pairing ‘hello’ with a wave or ‘yes’ with a head nod.
The system converts activity within seconds of the user’s intent and uses AI to turn electrical brain activity into on-screen text and to prompt a personalized animated avatar to move. It combines real-time operation with multimodal output in one pipeline rather than offline classification or a single output modality. Described as occurring ‘within seconds of the user’s intent’; a system-level proof of concept, with no quantitative accuracy or latency figures given in the text.
The two participants span different causes and residual abilities: one had impaired speech and movement after a brainstem stroke and was asked to silently speak five phrases and attempt to wave, nod, shake his hands and clap; the other had amyotrophic lateral sclerosis causing progressive speech paralysis and motor impairments, could still type on computer and phone at the time of the study, and was asked to vocalize ten phrases and imagine making ten gestures such as a shrug and thumbs up without moving his body except for facial muscles. It brings different injury types and different residual communication abilities into one decoding framework, suggesting the approach may extend to heterogeneous users. Two participants, with researcher-specified phrase and gesture sets; early feasibility evidence.
Perspective
The result is aimed at people with severe speech and movement impairment, such as after a brainstem stroke or with amyotrophic lateral sclerosis, validated in a controlled research setting with a surgically implanted 253-electrode cortical-surface array and researcher-specified phrase and gesture sets; it demonstrates the feasibility of one implant handling verbal and non-verbal communication together rather than a ready-to-use daily product.
The report does not provide specific decoding accuracy or latency values, training-data scale, follow-up duration or direct comparison with single-modality systems, nor does it address scalability beyond the specified gesture set; moreover, this is a news article rather than the full paper, and figures and supplementary materials were not loaded, so these quantitative questions remain open for a careful reader.
