Public articles linked to the same research event.
arXiv This work scales online reinforcement learning for multimodal deep-research agents to 128k context and 75+ tool-interaction turns and introduces Long-MDR, a three-component recipe of On-Policy Distillation Warmup, Progressive Horizon Expansion, and Entropy-Triggered Rescue, to address slow early reward growth and later entropy collapse in long-horizon training; at a 50-turn tool-interaction evaluation budget, the trained Long-MDR-9B ranks first on five of six benchmarks among the compared 7B–9B agents.
This work scales online reinforcement learning for multimodal deep-research agents to 128k context and 75+ tool-interaction turns and introduces Long-MDR, a three-component recipe of On-Policy Distillation Warmup, Progressive Horizon Expansion, and Entropy-Triggered Rescue, to address slow early reward growth and later entropy collapse in long-horizon training; at a 50-turn tool-interaction evaluation budget, the trained Long-MDR-9B ranks first on five of six benchmarks among the compared 7B–9B agents.
This work scales online reinforcement learning for multimodal deep-research agents to 128k context and 75+ tool-interaction turns and introduces Long-MDR, a three-component recipe of On-Policy Distillation Warmup, Progressive Horizon Expansion, and Entropy-Triggered Rescue, to address slow early reward growth and later entropy collapse in long-horizon training; at a 50-turn tool-interaction evaluation budget, the trained Long-MDR-9B ranks first on five of six benchmarks among the compared 7B–9B agents.
This work scales online reinforcement learning for multimodal deep-research agents to 128k context and 75+ tool-interaction turns and introduces Long-MDR, a three-component recipe of On-Policy Distillation Warmup, Progressive Horizon Expansion, and Entropy-Triggered Rescue, to address slow early reward growth and later entropy collapse in long-horizon training; at a 50-turn tool-interaction evaluation budget, the trained Long-MDR-9B ranks first on five of six benchmarks among the compared 7B–9B agents.