Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Long-MDR pushes online RL for multimodal deep-research agents to 128k context and 75 tool turns, with a 9B model leading five of six benchmarks among 7B–9B agents at a 50-turn evaluation budget

This work scales online reinforcement learning for multimodal deep-research agents to 128k context and 75+ tool-interaction turns and introduces Long-MDR, a three-component recipe of On-Policy Distillation Warmup, Progressive Horizon Expansion, and Entropy-Triggered Rescue, to address slow early reward growth and later entropy collapse in long-horizon training; at a 50-turn tool-interaction evaluation budget, the trained Long-MDR-9B ranks first on five of six benchmarks among the compared 7B–9B agents.