Long-MDR pushes online RL for multimodal deep-research agents to 128k context and 75 tool turns, with a 9B model leading five of six benchmarks among 7B–9B agents at a 50-turn evaluation budget
Related research and updatesSynopsis
This work scales online reinforcement learning for multimodal deep-research agents to 128k context and 75+ tool-interaction turns and introduces Long-MDR, a three-component recipe of On-Policy Distillation Warmup, Progressive Horizon Expansion, and Entropy-Triggered Rescue, to address slow early reward growth and later entropy collapse in long-horizon training; at a 50-turn tool-interaction evaluation budget, the trained Long-MDR-9B ranks first on five of six benchmarks among the compared 7B–9B agents.
[Uncaptioned image]
arXivInterpretation
The authors scale online multimodal deep-research RL training to 128k context and 75+ tool-interaction turns, which they state is the first online multimodal deep-research RL study trained at 128k context and the first with a 75 tool-turn horizon. Prior online RL for multimodal research agents has largely been confined to shorter contexts and interaction horizons; this work raises training to 128k context and 75 tool turns. The scale claim is stated by the authors and supported by a five-stage training configuration (tool budgets growing from 25 to 50 to 75 turns, 128k context cap) together with training-dynamics analysis.
Long-MDR consists of three components: On-Policy Distillation Warmup, Progressive Horizon Expansion, and Entropy-Triggered Rollback and Rescue, aimed at improving learning efficiency and stability in long-horizon RL. The authors report that directly extending a conventional RL configuration to this scale produces two behaviors: extremely slow early reward growth and later entropy collapse that plateaus reward before performance saturates; the three components target cold start, premature exposure to long trajectories, and stalled exploration respectively. The paper gives formal objectives for each component (including a reward-adaptive token-level distillation advantage, a stage-wise growing tool budget, and a collapse trigger based on training statistics rather than downstream benchmark scores) and reports trends in reward, validation performance, entropy, and tool usage during training.
At a 50-turn tool-interaction evaluation budget, Long-MDR-9B ranks first on five of six benchmarks among the compared 7B–9B agents. The result translates the long-horizon training recipe into leading benchmark performance within its size group, compared against 7B–9B multimodal research agents such as DeepVoyager-VL, SimpleSearch-VL, and POINTS-Seeker. The result comes from the benchmark comparison in Table 1, with a 50-turn tool-interaction cap; trajectories exceeding that cap are truncated and counted as incorrect. The table also lists proprietary direct-answer and agentic-workflow reference groups.
Training-dynamics analysis shows three phases—bootstrap, expansion, and rescue—and after entropy rescue the policy entropy recovers, reward rises again, and validation performance improves. The authors identify premature entropy collapse as a sustained entropy decrease accompanied by an outcome-reward plateau, and argue that policy convergence and entropy collapse are not the same event, motivating a state-dependent entropy intervention rather than a permanent entropy term. The paper reports that after rolling back to a pre-collapse checkpoint and resuming with a stronger entropy bonus and lower learning rate, entropy recovers, KL growth slows, and validation performance exceeds the original continuation; it also reports per-step rollout cost (14.8 minutes over steps 1–5, 25.3 minutes over steps 6–25, reaching 36.2 minutes at step 17) to convey the computational cost.
Perspective
This work targets online multimodal deep-research RL under a fixed model, tool environment, and maximum interaction budget, with training at 128k context and 75+ tool-interaction turns and evaluation at a 50-turn tool-interaction budget. It applies to training research agents that need long interaction histories, multi-round search, and visual evidence inspection, especially settings where rollouts are expensive and a single step can take tens of minutes. Its value lies in an actionable staged training procedure: use on-policy distillation warmup to establish basic research behavior, expand the tool budget as policy capability grows, and roll back and rescue when entropy decline coincides with a reward plateau.
Several equations, coefficients, and thresholds appear as placeholders in the provided text, so their exact values cannot be fully verified from it. Table 1 evaluates with a 50-turn tool-interaction cap while final training uses a 75-turn budget, and the BC-V3 validation value corresponds to different interaction budgets in different places, so cross-setting comparisons warrant care. The automated entropy-rescue trigger is described as an optional controller whose window, patience, and tolerances are parameters. In the appendix case study, tool-call counts differ between the table and the summary (75 numbered entries in the table versus 76 total calls reported in the summary), suggesting trace statistics still need a consistent accounting. These are scope and open questions rather than criticisms of the work.
