F-DSM factorizes waiting probability from the text vocabulary distribution, beating DSM on Japanese and English streaming recognition while cutting GPU memory
Related research and updatesSynopsis
The work first shows that the padding token can be removed from Delayed Streams Modeling (DSM) while maintaining competitive recognition performance, then proposes Factorized DSM (F-DSM), which separates the waiting probability from the distribution over the original LLM vocabulary, removing ASR-specific tokens from the text prediction space and allowing the large-vocabulary softmax to be skipped on waiting steps; experiments on the Corpus of Spontaneous Japanese and LibriSpeech show that F-DSM achieves better recognition performance than DSM, greatly reduces GPU memory use while maintaining similar training throughput, provides a small inference speed improvement through softmax skipping, and reduces the degradation in text-only perplexity observed with DSM.
Figure 1: DSM (left) jointly predicts waiting <p> , word-start <w> , and text tokens; F-DSM (right) separates waiting from text prediction over the original LLM vocabulary.
arXivInterpretation
The authors first verify that the padding token in DSM can be removed while still maintaining competitive recognition performance. Whereas DSM relies on the padding token and the word-start token being predicted together with normal text tokens using the same softmax, this result indicates the padding token is not required to sustain recognition performance. Supported by the recognition performance comparison reported in the abstract; specific numbers are not given in the loaded text.
They propose Factorized DSM (F-DSM), which separates the waiting probability from the distribution over the original LLM vocabulary. Unlike DSM, which predicts ASR-specific tokens together with normal text tokens using the same softmax, F-DSM removes ASR-specific tokens from the text prediction space and allows the large-vocabulary softmax to be skipped on waiting steps. The method description comes from the abstract and is a design-level contribution whose effects are supported by the subsequent experiments.
On the Corpus of Spontaneous Japanese and LibriSpeech, F-DSM achieves better recognition performance than DSM. This indicates the factorization yields not only efficiency gains but also an improvement in recognition quality relative to DSM. Evidence comes from experiments on two datasets; the abstract does not give specific error rates or improvement magnitudes.
F-DSM greatly reduces GPU memory use, maintains similar training throughput, provides a small inference speed improvement through softmax skipping, and reduces the degradation in text-only perplexity observed with DSM. These efficiency and language-modeling-preservation benefits are new engineering and modeling properties relative to DSM. The abstract reports directional conclusions across four metrics—memory, throughput, inference speed, and text-only perplexity—without providing specific values.
Perspective
The work targets LLM-based streaming automatic speech recognition, for systems that align acoustic and text streams on a common timeline. Its benefits are validated on the Corpus of Spontaneous Japanese and LibriSpeech, so the direct audience is streaming ASR researchers and engineering teams using DSM-style modeling and concerned with memory and inference efficiency. Methodologically, F-DSM's factorization design can be used by follow-up work to further separate ASR-specific control signals from the text prediction space.
The loaded text is abstract-level and does not include specific recognition error rates, memory reduction magnitudes, throughput or inference speed numbers, nor model scale, training data scale, or ablation settings. Readers should therefore still watch: how padding-token removal behaves with larger vocabularies or more languages, how the factorization affects scenarios with a high proportion of waiting steps, and the concrete magnitude of the text-only perplexity improvement. These are open questions for follow-up verification.
