Public articles linked to the same research event.
arXiv The work proposes LexiHorizon, a long-horizon deep-search training framework that expands the trajectory context budget to 128K, manages accumulated retrieval content with a context window that keeps only recent tool observations in full while preserving the reasoning history, and adds an outcome-gated search-effort reward that applies only to trajectories with nonzero answer reward, training Qwen3.5-9B online with GSPO; the resulting LexiHorizon-9B outperforms both its base model and MiroThinker-1.7-mini on XBench, WebWalkerQA, and BrowseComp-ZH, with maximum absolute gains of 8.7 and 23.8 percentage points, and ablations show that a 64K budget lowers the valid answer rate from 99.88% to 66.97% while removing the window reduces accuracy on all three benchmarks.
The work proposes LexiHorizon, a long-horizon deep-search training framework that expands the trajectory context budget to 128K, manages accumulated retrieval content with a context window that keeps only recent tool observations in full while preserving the reasoning history, and adds an outcome-gated search-effort reward that applies only to trajectories with nonzero answer reward, training Qwen3.5-9B online with GSPO; the resulting LexiHorizon-9B outperforms both its base model and MiroThinker-1.7-mini on XBench, WebWalkerQA, and BrowseComp-ZH, with maximum absolute gains of 8.7 and 23.8 percentage points, and ablations show that a 64K budget lowers the valid answer rate from 99.88% to 66.97% while removing the window reduces accuracy on all three benchmarks.
The work proposes LexiHorizon, a long-horizon deep-search training framework that expands the trajectory context budget to 128K, manages accumulated retrieval content with a context window that keeps only recent tool observations in full while preserving the reasoning history, and adds an outcome-gated search-effort reward that applies only to trajectories with nonzero answer reward, training Qwen3.5-9B online with GSPO; the resulting LexiHorizon-9B outperforms both its base model and MiroThinker-1.7-mini on XBench, WebWalkerQA, and BrowseComp-ZH, with maximum absolute gains of 8.7 and 23.8 percentage points, and ablations show that a 64K budget lowers the valid answer rate from 99.88% to 66.97% while removing the window reduces accuracy on all three benchmarks.
The work proposes LexiHorizon, a long-horizon deep-search training framework that expands the trajectory context budget to 128K, manages accumulated retrieval content with a context window that keeps only recent tool observations in full while preserving the reasoning history, and adds an outcome-gated search-effort reward that applies only to trajectories with nonzero answer reward, training Qwen3.5-9B online with GSPO; the resulting LexiHorizon-9B outperforms both its base model and MiroThinker-1.7-mini on XBench, WebWalkerQA, and BrowseComp-ZH, with maximum absolute gains of 8.7 and 23.8 percentage points, and ablations show that a 64K budget lowers the valid answer rate from 99.88% to 66.97% while removing the window reduces accuracy on all three benchmarks.