Skip to main content
Back to timeline
arXivSource publication:

AdaLoop lets audio language models spend reasoning depth by question difficulty, lifting average accuracy 2.9 to 3.8 points on three backbones

Synopsis

AdaLoop replaces the linear projection between an audio encoder and a language model with a lightweight recurrent module in which a shared, question-guided transformer block iteratively refines the audio representation and a learned halting mechanism decides how many steps each audio–question pair needs; across Qwen2.5-Omni-7B, Kimi-Audio-7B-Instruct, and MiMo-Audio-7B-Instruct on MMSU, MMAU-Pro, and MMAR it raises average accuracy by 2.9 to 3.8 points, with gains concentrated on perception-heavy subtasks, and the learned depth distribution shows signal-level questions averaging 5.7 iterations versus 2.6 for semantic ones.

Source-provided article image: AdaLoop: Adaptive-Depth Latent Reasoning for Audio Language Models
Fig. 1 ·

Fig. 1. Overview of AdaLoop. The Recurrent Latent Block (dashed) applies question-guided cross-attention, scale-aware self-attention with four window sizes (w), and a feed-forward update. Shared parameters iterate K times; a halting unit conditioned on both the pooled audio state and the question decides when to stop.

arXiv · Page 2

Interpretation

AdaLoop swaps the linear projection connecting the audio encoder to the language model for an adaptive-depth recurrent module: a parameter-shared Recurrent Latent Block refines the audio representation through question-guided cross-attention and scale-aware self-attention, while a halting unit emits a scalar halting probability after each iteration conditioned on both the current state and the question; iteration stops once the cumulative halting signal crosses the threshold, and the output is a weighted mixture of the per-iteration representations. Prior audio iterative-refinement methods such as RecurTrace, AURAL, Thinking-in-Depth, and CoRELoop each target a single task with a fixed number of iterations, and generalist audio language models run every query at one fixed depth; AdaLoop conditions depth on both audio and question, so the same model uses one or two passes for language identification but six or more for a fine pitch comparison. The method is specified with full equations for question-guided cross-attention, scale-aware self-attention whose grouped windows double in receptive field, the halting probability, and the weighted-mixture output, and the module is trainable end to end; the loss adds a ponder cost to the standard next-token objective to encourage fewer iterations.

Across three architecturally distinct audio language models, AdaLoop attains the highest average and the highest score on every individual benchmark among MMSU, MMAU-Pro, and MMAR: average gains over Base of 3.8 points on Qwen2.5-Omni, 3.4 on Kimi-Audio, and 2.9 on MiMo-Audio, with all twelve subcategories improving on Qwen2.5-Omni and the largest single jump on MMSU perception at +8.6 points. Ablations separate the sources of gain: fine-tuning only the original connector on the same 50K set (SFT) yields 1.2 points, switching to the recurrent architecture at fixed depth four (Fixed-4) adds 0.9, and adaptive depth adds a further 1.7, so the largest step comes from the adaptive mechanism rather than the training data or the recurrent architecture alone. Table results order consistently across three backbones and three benchmarks; the ablation table shows the largest drop when adaptive halting is replaced by fixed depth four, 1.6 points lost when the question is removed from the halting input, 1.2 points from scale-aware attention, and 0.8 points each from removing the ponder cost or capping at four iterations.

Learned reasoning depth tracks task nature: on Qwen2.5-Omni, signal-level questions such as pitch comparison, loudness ordering, and rate estimation receive 5.7 iterations on average, nearly triple the 2.6 used for semantic questions; perceptual tasks sit at 4.3 and cultural reasoning at 3.1, with an overall average of 3.8 iterations ranging from a single pass (9.3% of items) to the maximum eight (4.2%). The allocation emerges from the ponder cost alone with no explicit depth labels, showing the model discovers that signal-level judgments benefit from repeated attention at different scales while semantic questions can be resolved from a single high-level embedding; distributions also shift with backbone strength, with Kimi-Audio averaging 4.2 iterations and MiMo-Audio 3.9, reaching 5.4 on signal-level tasks. Depth distributions are broken down by MMAR category and reported across backbones; the ablation indicates eight iterations provide headroom the model actually uses, since the few questions reaching eight passes include the hardest signal-level comparisons.

The module is lightweight and architecture-agnostic: the RLB matches the original connector's hidden dimension, uses 8 attention heads split into 4 scale groups and one feed-forward layer of expansion ratio 4, adds 2.4% to 2.9% of the base model's parameters, and plugs between any audio encoder and language backbone without modifying either, with the audio encoder frozen throughout. Unlike methods that modify encoders or language backbones or are designed for a single task, AdaLoop offers transferable control of reasoning depth at under 3% parameter overhead, and the authors note that recent models including Audio Flamingo and MiniCPM-o could also benefit. Training runs on 4 NVIDIA A100-80G GPUs with per-GPU batch size 2 and gradient accumulation over 4 steps (effective batch 32); two-stage training first freezes the encoder and language model to train the RLB and halting unit, then jointly fine-tunes the RLB, halting unit, and language model; audio input is capped at 3000 tokens (about 30 seconds).

Perspective

The result targets generalist audio language models that need fine-grained acoustic judgment at inference time: AdaLoop is inserted between the audio encoder and the language backbone without modifying either, adds under 3% of parameters, and operates on audio capped at 3000 tokens (about 30 seconds). It applies to verifiable question–answer families covering speech prosody, speakers and dialogue, speech content, sound events, music, and multi-clip scenes, trained on 50,000 audio–question pairs drawn from the verifiable audio QA families of EvoAudio, with source audio from LibriSpeech, FSD50K, AudioSet, AudioCaps, and MELD plus synthesis by Qwen3-TTS, LeVo, and procedural mixing via Scaper. The authors name combining adaptive depth with richer training curricula and parameter-efficient adaptation as the natural next step.

A careful reader would still watch how stable the depth allocation is when moving to other encoder–language-model pairings, since the depth distributions and ablations come from specific backbones such as Qwen2.5-Omni; the training data are all verifiable synthetic or labeled QA families, so depth behavior on open-domain real audio is not shown; audio is uniformly truncated to about 30 seconds, leaving open how iteration counts and gains change for longer recordings; and MiMo-Audio reports only Base and AdaLoop, so the relative contributions of data and architecture are less clear there than for the other two backbones.

Sources