Skip to main content
Back to timeline
arXivSource publication:

Reinforcement learning teaches a language model to search for its own evidence, beating Claude Opus 4.5 on 265 questions at about 5% inference cost

Related research and updates

Synopsis

The work builds an agentic forecasting environment, dataset, and harness from 2,100+ resolved Polymarket questions with layered leak filtering, letting the agent acquire its own context at rollout time via web search, page reading, and financial time series, and trains Qwen3.5-35B-A3B (3B active parameters) with single-epoch GRPO under a Brier-score reward; training improves calibration by 30-40% and cuts search attempts from 3.8 to 2.25 per rollout, and in an identical harness the trained policy finishes ahead of every frontier model tested at evidence-based forecasting, including Claude Opus 4.5 (soft-Brier 0.254 vs. 0.256, n=265), at about 5% of the inference cost, with the widest margin on the hardest questions the crowd itself had not decided.

Source-provided article image: Do Your Own Research: Learning to Forecast by Learning to Search
Figure 1 ·

Figure 1: Soft-Brier on the full held-out test split ( n = 265 n{=}265 ) and on the pre-declared uncertain subset (cutoff price in [ 0.30 , 0.70 ] [0.30,0.70] , n = 104 n{=}104 ). The range [ 0 , 0.24 ] [0,0.24] is compressed for readability.

arXiv

Interpretation

It brings the skill of gathering evidence inside the reward: the agent acquires its own context at rollout time rather than freezing research context before training or deploying agentic research only at test time. Prior forecasting work either froze research context or used agentic research only at test time, so evidence gathering was never shaped by the reward; here search behavior itself becomes a trainable object. An environment, dataset, and harness built from 2,100+ resolved Polymarket questions, with layered leak filtering restricting usable information to what was published before each question's cutoff.

Training changes how the agent interacts with information: calibration improves 30-40%, and search attempts fall from 3.8 to 2.25 per rollout. The gain is not merely more retrieval but the formation of evidence discipline, obtaining better-calibrated probabilities with fewer searches. Qwen3.5-35B-A3B (3B active parameters) trained with single-epoch GRPO under a Brier-score reward, with reported changes in calibration and search counts.

In an identical harness, the trained policy finishes ahead of every frontier model tested at evidence-based forecasting, including Claude Opus 4.5 (soft-Brier 0.254 vs. 0.256, n=265), at about 5% of the inference cost. It places a small-active-parameter model with trainable retrieval against frontier models under the same evaluation conditions and reports an order-of-magnitude cost difference. A controlled comparison against four frontier models within the same harness, with n=265 and soft-Brier as the metric.

Its margin is widest on the hardest questions, the ones the crowd itself had not decided, and the environment, dataset, and per-rollout records are released as a reusable harness. It locates the forecasting gain where information is thinnest and consensus weakest, and releases reusable assets for follow-up temporal forecasting agent research. Reported performance differences stratified by question difficulty, plus a stated release of the environment, dataset, and per-rollout records.

Perspective

The result targets real-world event probability forecasting under information cutoffs, where the agent must gather its own evidence, and applies to settings with access to web search, page reading, and financial time series. Its reusable artifacts, the environment, dataset, and per-rollout records, provide a basis for training and same-condition comparison of temporal forecasting agents, especially for studying how retrieval behavior is shaped by reward and how low-cost models perform on difficult questions. The evaluation conclusion is scoped to comparison with the four frontier models tested within the same harness and to the n=265 soft-Brier comparison.

The visible text is an abstract and does not show the specific question-selection criteria, the implementation of layered leak filtering, GRPO hyperparameters, how the 30-40% calibration improvement is measured, or the individual results for the four frontier models. The soft-Brier gap of 0.254 vs. 0.256 is small, and its stability, confidence intervals, and behavior across question subsets remain to be confirmed in the paper. How the hardest questions are defined and how crowd indecision is determined are also questions a reader would continue to watch.

Sources