Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

F-DSM factorizes waiting probability from the text vocabulary distribution, beating DSM on Japanese and English streaming recognition while cutting GPU memory

The work first shows that the padding token can be removed from Delayed Streams Modeling (DSM) while maintaining competitive recognition performance, then proposes Factorized DSM (F-DSM), which separates the waiting probability from the distribution over the original LLM vocabulary, removing ASR-specific tokens from the text prediction space and allowing the large-vocabulary softmax to be skipped on waiting steps; experiments on the Corpus of Spontaneous Japanese and LibriSpeech show that F-DSM achieves better recognition performance than DSM, greatly reduces GPU memory use while maintaining similar training throughput, provides a small inference speed improvement through softmax skipping, and reduces the degradation in text-only perplexity observed with DSM.