Text-only post-training lifts Qwen2.5-Omni-7B's geometric mean of nine reasoning scores by 25.83% with 56.6% fewer GPU-hours
Related research and updatesSynopsis
Diagnosing multi-hop reasoning difficulties and partial decoupling in the local optimization of perception and reasoning objectives in omni-modal LLMs, this work proposes a text-centric post-training paradigm in which text-only training carries the main reasoning optimization and reduced-data native audio-visual RL then refines perception; the best text-only configuration (supervised fine-tuning followed by RL) raises Qwen2.5-Omni-7B's geometric mean of nine reasoning scores by 25.83% over the base model with 56.6% fewer GPU-hours than the complete native audio-visual route, training on data synthesized entirely by a text-only LLM raises this geometric mean by 21.01%, and refinement restores perception above the base level while retaining 93.5% of the reasoning gain.
Figure 1: (a) A diagnostic case of multi-hop reasoning failure despite correct answers to corresponding single-hop questions. (b) Text-only training provides the main reasoning optimization, followed by reduced-data native audio-visual RL to restore perception while retaining most reasoning gains.
arXivInterpretation
Diagnostics reveal multi-hop reasoning difficulties even when all corresponding single-hop questions are answered correctly, and suggest partial decoupling in the local optimization of perception and reasoning objectives. Whereas improving joint audio-visual reasoning typically incurs substantial data construction and training costs, this work locates the bottleneck in multi-hop reasoning and in how the two objectives are optimized, informing how post-training should divide labor. Based on the diagnostics described by the authors; the abstract does not report the specific datasets, sample sizes, or statistics used.
Text-only reasoning training yields gains across data sources, model scales, and families; the best text-only configuration (supervised fine-tuning followed by RL) raises Qwen2.5-Omni-7B's geometric mean of nine reasoning scores by 25.83% over the base model while using 56.6% fewer GPU-hours than the complete native audio-visual route. Moves reasoning gains away from the expensive audio-visual data route to text-only training, and reports generality across data sources, scales, and families alongside the compute comparison. Uses the geometric mean of nine reasoning scores as the metric, compared against the base model and the complete native audio-visual route, with a reported GPU-hour difference; the abstract does not list the nine scores.
Training on data synthesized entirely by a text-only LLM, with no audio-visual data in construction or training, still raises this geometric mean by 21.01%. Shows the reasoning gain does not depend on real audio-visual data for construction or training, reducing reliance on multimodal data pipelines. A geometric-mean gain for a single configuration; the abstract reports no variance, repetitions, or statistical tests.
Because text-only training degrades perception, the authors propose a text-centric post-training paradigm: text-only training provides the main reasoning optimization, and reduced-data native audio-visual RL then refines perception; refinement uses about 90% fewer input tokens, restores perception above the base level, and retains 93.5% of the best-performing text-only pipeline's reasoning gain. Explicitly divides reasoning and perception optimization, using a small amount of audio-visual RL to repair the perception loss from text-only training while keeping most of the reasoning gain and cutting input-token cost. Reports the input-token reduction, the direction of perception change relative to the base, and the retained reasoning-gain fraction; the abstract does not give the specific perception metrics or datasets.
Perspective
The paradigm targets research and engineering teams that need to improve joint audio-visual reasoning in omni-modal models under limited compute and data, in settings where reasoning scores are the main objective and a small amount of native audio-visual RL can be spent to repair perception. Its conclusions rest on post-training experiments on models such as Qwen2.5-Omni-7B, and the abstract reports gains across data sources, model scales, and families, so it can inform the design of post-training pipelines for comparable omni-modal models; the text-only main training plus reduced-data audio-visual refinement configuration suits pre-deployment tuning with a constrained data-construction budget that must still preserve perception.
Readers should still watch how the multi-hop reasoning difficulty and the partial perception-reasoning decoupling were measured and how far they generalize; which capabilities the nine reasoning scores cover and whether the geometric-mean gain is consistent across individual scores; which perception metrics and data underlie the claim that perception is restored above the base level; and how the text-only synthetic-data route behaves at larger scales and with more modalities. Because only the abstract is available here, figures and experimental details are not included, and answers to these questions should be checked against the original text.
