DeepFisFis processes 5 ms audio segments in about 2.5 ms, detecting mouse ultrasonic vocalizations in real time and triggering closed-loop stimulation
Synopsis
The work introduces DeepFisFis, a waveform-based neural network that classifies consecutive 5 ms audio segments to detect mouse ultrasonic vocalizations (USVs) while they are being produced, processing each segment in approximately 2.5 ms and thus faster than the incoming audio stream; in a deployed closed-loop system, detections triggered an external stimulus, demonstrating online control of ongoing vocal behaviour, and event-triggered acquisition preserved more than 99% of vocalization time while retaining only approximately 22% of the continuous recording.
Interpretation
DeepFisFis detects USVs while they are being produced, rather than analysing them only after acquisition as is typical. Where USVs have usually served as a retrospective readout after acquisition, this work moves detection into the ongoing vocalization, making the call itself a signal usable in real time. The abstract reports that the network classifies consecutive 5 ms segments directly from the waveform and describes this as high accuracy; the specific accuracy value is not given in the text.
The network processes each segment in approximately 2.5 ms, faster than the 5 ms audio segment stream, which is what makes real-time operation possible. Processing faster than the audio arrives is the key condition that lets detections guide ongoing experimental interventions. The abstract provides the comparison between approximately 2.5 ms per segment and the 5 ms segment length, which is evidence at the level of the time budget.
In a deployed closed-loop system, detections triggered an external stimulus, demonstrating online control of ongoing vocal behaviour. This moves detection from offline analysis toward closed-loop intervention, allowing USVs to be used for causal experimental manipulation rather than behavioural description alone. The abstract states that the system was deployed and that stimulation was triggered; the text does not give the stimulus type, number of trials, or magnitude of behavioural change.
Event-triggered acquisition preserved more than 99% of vocalization time while retaining only approximately 22% of the continuous recording. This provides a quantitative basis for selective storage gated around detected calls, substantially reducing data volume without appreciable loss of vocalization content. The abstract reports two percentage figures (more than 99% and approximately 22%) but does not describe the dataset size or the statistical treatment behind them.
Perspective
The result is aimed at experimental settings that require reacting while a mouse is vocalizing, such as a closed-loop system in which detections trigger an external stimulus, and at event-triggered acquisition that stores data around detected calls. For researchers who want to reduce continuous-recording storage while preserving vocalization content, the abstract's contrast between more than 99% of vocalization time and approximately 22% of the continuous recording offers a direct reference; for those who want to feed a behavioural readout into real-time intervention, the approximately 2.5 ms per segment, faster than the 5 ms segment, is the precondition for usability.
The reading scope here is incomplete, covering only the abstract and the competing-interest statement; the main text, figures, and supplementary materials were not included, so the "high accuracy" mentioned in the abstract has no specific value, and the stimulus parameters, trial counts, and behavioural effect sizes of the closed-loop experiment cannot be checked. The dataset size, animal numbers, and statistical methods behind the more than 99% and approximately 22% figures also do not appear in the text. Readers who want to judge stability across strains, noise environments, or longer recordings would still need the full paper.
