STAM-ASR extends a pretrained AudioLLM with speaker-temporal anchoring and memory for multi-speaker ASR, evaluated on AMI, ICSI, LibriCSS, and NOTSOFAR-1
Synopsis
STAM-ASR is a lightweight framework that, without relying on an external diarization system and without explicit speech separation, learns speaker activity and speaker-aware representations directly from intermediate AudioLLM features to give explicit who-and-when cues that modulate the AudioLLM's semantic representation, and maintains fixed-size speaker and conversational memories to carry complementary context across turns; it is evaluated on AMI, ICSI, LibriCSS, and NOTSOFAR-1 across close-talk, far-field, overlapping, and cross-domain conditions, with reported results showing that speaker-temporal conditioning and memory provide complementary benefits while the gap between reference and predicted speaker activity identifies robust speaker tracking as a key remaining challenge.
Figure 1: Overview of STAM-ASR with internal diarization, speaker-temporal modulation, memory construction.
arXivInterpretation
STAM-ASR proposes a lightweight framework that adds speaker-temporal anchoring and memory mechanisms on top of an already pretrained AudioLLM for multi-speaker ASR. Unlike the common practice of relying on an external diarization system, the framework learns speaker activity and speaker-aware representations directly from intermediate AudioLLM features. The source presents this as a framework description and design motivation, without parameter counts, training data scale, or ablation details; the evidence is the authors' statement of the method.
The framework injects explicit who and when cues into the AudioLLM's semantic representation without performing explicit speech separation. Speaker attribution information is used as a conditioning signal that modulates the semantic representation, rather than separating speech first and then recognizing it, avoiding dependence on a separation module. Stated in the source as a method description, without item-by-item numerical comparison against separation-based baselines.
STAM-ASR maintains fixed-size speaker and conversational memories to carry complementary context across turns. For long conversations with turn-taking, overlap, and speakers reappearing over time, fixed-capacity memory is introduced to preserve cross-turn information. The source states that the memories are fixed-size but does not give memory capacity, update rules, or a quantitative analysis of their effect on performance.
Evaluation on four datasets, AMI, ICSI, LibriCSS, and NOTSOFAR-1, across close-talk, far-field, overlapping, and cross-domain conditions reports that speaker-temporal conditioning and memory provide complementary benefits. The evaluation spans multiple acoustic and domain conditions and points to the gap between reference and predicted speaker activity, marking robust speaker tracking as a key remaining challenge. The source summarizes the conclusion as reported results and does not give specific word error rate or speaker attribution metric values in the abstract.
Perspective
The work targets multi-speaker ASR and is suited to teams that already have a pretrained AudioLLM, evaluated under close-talk, far-field, overlapping, and cross-domain conditions on AMI, ICSI, LibriCSS, and NOTSOFAR-1. It enables follow-up work to plug speaker-temporal cues and cross-turn memory directly into an AudioLLM without an external diarization system and without explicit speech separation, and it is relevant to applications such as meeting transcription that need speaker attribution over long conversations.
What is available here is abstract-level information, lacking specific word error rates, speaker attribution metrics, memory capacity settings, and ablation results, so the magnitude of the separate contributions of conditioning and memory cannot be judged. The source notes a gap between reference and predicted speaker activity and identifies robust speaker tracking as a key remaining challenge, suggesting performance may be limited under heavy overlap or frequent speaker reappearance. The abstract also does not state dependence on AudioLLM scale, language, or domain, which are open questions a reader should watch when assessing transferability.
