Skip to main content
Back to timeline
arXivSource publication:

AURAL reaches CoT-RL-comparable performance on speech language models with 11.8x faster time to first answer token via latent reasoning and joint chunk prediction

Synopsis

The work introduces AURAL, which models a distribution over multiple plausible reasoning continuations in latent space and jointly predicts chunks of future states, and with AuralReason-683K (683K bilingual speech utterances, about 1,000 hours) for initial supervision plus AURAL-RL, it achieves performance comparable to CoT-RL across two backbones with larger gains over the respective supervised checkpoints on most metrics, and on Qwen2.5-Omni reduces time to the first answer token by 11.8x, from 1.22 to 0.10 s, versus 0.05 s for direct answering.

Source-provided article image: AURAL: Adaptive Latent Reasoning with Joint Chunk for Speech Language Models
Figure 1 ·

Figure 1: Overview of AURAL. Top left: AURAL-SFT pools CoT embeddings into latent targets under compression factor c c , then predicts a Gaussian mixture over a joint chunk of m m future states that reenters the backbone in one pass, while the language head keeps each state readable and emits EOL . Bottom left: AURAL-RL scores 8 rollouts per prompt with a quality-gated conciseness reward, making latent depth adaptive. Right: AURAL-RL is comparable to CoT-RL over eleven metrics while cutting time to first token from 1.22 1.22 to 0.10 0.10 s ( 11.8 × 11.8\times ).

arXiv

Interpretation

AURAL models a distribution over multiple plausible reasoning continuations in latent space and jointly predicts chunks of future states, reducing sequential forward passes and reasoning latency. Relative to existing latent reasoning methods limited by single-path supervision and reasoning budgets that do not adapt to problem difficulty, this design adds multi-path distribution modeling and joint chunk prediction. Method description at the abstract level; no ablation or chunk-size details are given.

It constructs AuralReason-683K: 683K bilingual speech utterances (about 1,000 hours) with concise CoT for emotion recognition, empathetic dialogue, and general reasoning, providing initial supervision for latent reasoning. It supplies large-scale bilingual speech supervision with concise CoT for latent reasoning, rather than relying only on explicit CoT text. The abstract reports data scale and task coverage but not construction pipeline or quality evaluation details.

AURAL-RL explores beyond these traces, rewarding concise reasoning that yields high-quality answers and adapting reasoning effort to each problem. The reward signal ties conciseness to answer quality so that reasoning budget varies with problem difficulty rather than being fixed. Training-objective description at the abstract level; reward form and hyperparameters are not given.

Across two backbones, AURAL-RL achieves performance comparable to CoT-RL with larger gains over the respective supervised checkpoints on most metrics, and analysis shows harder questions elicit more latent reasoning steps. It shows difficulty-adaptive reasoning-step behavior while maintaining performance comparable to explicit CoT reinforcement learning. The abstract reports comparisons on two backbones and a difficulty-versus-steps analysis, without listing specific metric values or statistical tests.

Perspective

The result targets speech language model interaction settings that need low-latency responses, especially emotion recognition, empathetic dialogue, and general reasoning tasks; it is relevant to researchers and engineering teams seeking to compress time to first answer token while preserving reasoning quality. The method is validated on the two backbones described in the abstract and is paired with the described 683K bilingual speech supervision data, suited to settings with speech input and reasoning-style answers.

Based only on the abstract, specific evaluation metric values, statistical significance, chunk size, and reward function details are not visible, so the magnitude of 'comparable' performance and the stability of the latency gain still need confirmation in the full text; the generality of the difficulty-linked increase in latent reasoning steps across tasks and backbones, and the practical experiential impact of the gap between 0.10 s and 0.05 s for direct answering, are also worth watching.

Sources