Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Text-only post-training lifts Qwen2.5-Omni-7B's geometric mean of nine reasoning scores by 25.83% with 56.6% fewer GPU-hours

Diagnosing multi-hop reasoning difficulties and partial decoupling in the local optimization of perception and reasoning objectives in omni-modal LLMs, this work proposes a text-centric post-training paradigm in which text-only training carries the main reasoning optimization and reduced-data native audio-visual RL then refines perception; the best text-only configuration (supervised fine-tuning followed by RL) raises Qwen2.5-Omni-7B's geometric mean of nine reasoning scores by 25.83% over the base model with 56.6% fewer GPU-hours than the complete native audio-visual route, training on data synthesized entirely by a text-only LLM raises this geometric mean by 21.01%, and refinement restores perception above the base level while retaining 93.5% of the reasoning gain.