STAM-ASR extends a pretrained AudioLLM with speaker-temporal anchoring and memory, evaluated on AMI, ICSI, LibriCSS, and NOTSOFAR-1 for multi-speaker ASR
Synopsis
The work proposes STAM-ASR, a lightweight framework that extends an already pretrained AudioLLM for multi-speaker ASR without relying on an external diarization system, learning speaker activity and speaker-aware representations directly from intermediate AudioLLM features to provide explicit who-and-when cues that modulate the semantic representation, while maintaining fixed-size speaker and conversational memories to carry complementary context across turns; it is evaluated on AMI, ICSI, LibriCSS, and NOTSOFAR-1 across close-talk, far-field, overlapping, and cross-domain conditions, with reported results showing that speaker-temporal conditioning and memory provide complementary benefits and that the gap between reference and predicted speaker activity identifies robust speaker tracking
Figure 1: Overview of STAM-ASR with internal diarization, speaker-temporal modulation, memory construction.
arXivInterpretation
STAM-ASR learns speaker activity and speaker-aware representations directly from intermediate AudioLLM features without an external diarization system, providing explicit who-and-when cues to modulate the semantic representation. Compared with common pipelines that depend on external diarization or explicit speech separation, this framework internalizes speaker-temporal cues into feature modulation of the AudioLLM without explicit speech separation. The abstract states the framework extends an already pretrained AudioLLM and is evaluated on AMI, ICSI, LibriCSS, and NOTSOFAR-1 across close-talk, far-field, overlapping, and cross-domain conditions; specific numbers are not given in the loaded text.
STAM-ASR maintains fixed-size speaker and conversational memories to carry complementary context across turns. Beyond speaker-temporal conditioning, the memory mechanism lets the model use cross-turn information rather than only the current segment. The abstract reports that speaker-temporal conditioning and memory provide complementary benefits, but no ablation numbers or memory capacity settings are given.
Evaluation on AMI, ICSI, LibriCSS, and NOTSOFAR-1 shows that the gap between reference and predicted speaker activity identifies robust speaker tracking as a key remaining challenge. The work extends evaluation to close-talk, far-field, overlapping, and cross-domain conditions and explicitly points to speaker tracking, rather than recognition alone, as the remaining bottleneck. The conclusion is based on evaluation across four datasets, but the loaded text does not provide specific error rates or comparison figures.
Perspective
The work targets research and engineering settings that extend an already pretrained AudioLLM for multi-speaker ASR, applicable to close-talk, far-field, overlapping, and cross-domain conditions; its value lies in offering a lightweight framework that does not rely on an external diarization system and in explicitly naming robust speaker tracking as a follow-up direction. For readers seeking to improve multi-speaker recognition without introducing an external diarization module, this approach is directly relevant.
The loaded text contains only the abstract and gives no specific recognition error rates, speaker attribution accuracy, memory capacity, training data scale, or numerical comparison with baselines, so the magnitude of gains cannot be judged. The gap between reference and predicted speaker activity is identified as a key challenge, but the text does not describe how that gap varies across datasets or conditions. Readers who need to assess applicability to their own setting would still need the experimental tables and implementation details in the original.
