Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Reinforcement learning teaches a language model to search for its own evidence, beating Claude Opus 4.5 on 265 questions at about 5% inference cost

The work builds an agentic forecasting environment, dataset, and harness from 2,100+ resolved Polymarket questions with layered leak filtering, letting the agent acquire its own context at rollout time via web search, page reading, and financial time series, and trains Qwen3.5-35B-A3B (3B active parameters) with single-epoch GRPO under a Brier-score reward; training improves calibration by 30-40% and cuts search attempts from 3.8 to 2.25 per rollout, and in an identical harness the trained policy finishes ahead of every frontier model tested at evidence-based forecasting, including Claude Opus 4.5 (soft-Brier 0.254 vs. 0.256, n=265), at about 5% of the inference cost, with the widest margin on the hardest questions the crowd itself had not decided.