LexiHorizon trains a 9B search agent with a 128K context budget and a reasoning-preserving observation window, beating its base model and MiroThinker-1.7-mini on XBench, WebWalkerQA, and BrowseComp-ZH
Related research and updatesSynopsis
The work proposes LexiHorizon, a long-horizon deep-search training framework that expands the trajectory context budget to 128K, manages accumulated retrieval content with a context window that keeps only recent tool observations in full while preserving the reasoning history, and adds an outcome-gated search-effort reward that applies only to trajectories with nonzero answer reward, training Qwen3.5-9B online with GSPO; the resulting LexiHorizon-9B outperforms both its base model and MiroThinker-1.7-mini on XBench, WebWalkerQA, and BrowseComp-ZH, with maximum absolute gains of 8.7 and 23.8 percentage points, and ablations show that a 64K budget lowers the valid answer rate from 99.88% to 66.97% while removing the window reduces accuracy on all three benchmarks.
(a) Response length.
arXivInterpretation
Long-horizon search reinforcement learning is framed as a training stability problem arising from premature context exhaustion, interference from accumulated retrieval content, and weak incentives for sustained exploration. Prior deep-search work often assumes relatively stable retrieval and short-horizon tool interaction, or focuses on question decomposition, training data, and trajectory construction; this work makes stability under long interaction horizons the training target itself. Supported by problem formulation and controlled ablations on three benchmarks: under a 64K budget the valid answer rate falls to 66.97% while reward and actor entropy both decline toward zero.
A reasoning-preserving context window keeps the complete reasoning and tool-action history while retaining full observations only for the most recent tool interactions, replacing earlier observations with markers. Unlike trajectory summarization or compression, this design discards only raw tool outputs while leaving information already written into the reasoning history available to the policy. Across evaluation trajectories the window is activated in more than 80% of cases and reduces overall active context by 79.1%; of 55 trajectories whose removed observations contained valid answer evidence, the reasoning history retained it in 44, an 80% retention rate, with all 11 failures in trajectories exceeding 40 tool calls.
An outcome-gated search-effort reward gives a bounded bonus for tool calls that saturates after a set number of calls and is multiplicatively gated by answer reward, so only trajectories with nonzero answer reward receive it. Unlike unconditionally rewarding tool use or adding denser process rewards, this design ties the exploration incentive to answer quality and prevents zero-answer-reward trajectories from gaining by making more calls. The reward combines an LLM-judged ordinal correctness score, the gated bounded tool bonus, and format and repetition penalties, with the policy optimized by GSPO over online rollouts with live tool interaction.
The resulting LexiHorizon-9B ranks first or second on all three benchmarks while also improving reliability and efficiency. It improves over MiroThinker-1.7-mini by 23.8, 4.6, and 6.5 percentage points on XBench, WebWalkerQA, and BrowseComp-ZH, and outperforms the larger Qwen3.5-27B on XBench and BrowseComp-ZH. Tool call failure rate drops from 12.80% to 2.00%, valid answer rate rises from 87.60% to 99.88%, average policy-generated tokens fall from 9,947 to 6,124, and end-to-end throughput reaches 1,069 tokens per second.
Perspective
The framework targets long-horizon deep search where retrieval is sensitive to query wording and repeated query reformulation is required, and it applies to training search agents that interleave ReAct-style reasoning with tool calls and can be optimized from outcome rewards. It lets a 9B policy sustain query exploration under an extended context budget while keeping active context small: when the window is activated, overall active context drops by 79.1%, and after training the policy averages 54.9 tool calls versus 29.5, while the 95th-percentile active context length falls from 120,818.9 to 28,324.5 tokens. For teams training long-horizon search agents under limited memory and generation budgets, this combination offers reusable context and reward design; trajectory reset as an auxiliary mechanism improves BrowseComp-ZH by 1.9 percentage points, leaves WebWalkerQA unchanged, and decreases XBench by 2.3 percentage points, making it better suited as a recovery mechanism for difficult samples.
Worth watching: evidence retention falls to 65.62% in the longest trajectories with more than 40 tool calls, so information preservation in very long trajectories still has room; answer correctness comes from an LLM-judged ordinal score, so judge consistency affects the reward signal; the benefit of trajectory reset is task-dependent, and the sensitivity of its trigger threshold and reset cap is not expanded in the main text; in addition, training dynamics and length distributions are cited by figure number, so exact curve values require consulting the original figures.
