Skip to main content
Back to timeline
arXivSource publication:

Forecast-Dojo builds a replayable forecasting environment from 1,568 Polymarket events and 18.8M dated news articles, where research tools lower Brier score for all 12 models yet every model still trails market forecasts

Synopsis

The authors introduce Forecast-Dojo, a replayable environment that combines resolved prediction-market questions with dated news so LLM forecasting agents can research an event and revisit their predictions at successive historical dates; it contains 1,568 Polymarket events split by time into training and evaluation periods and 18.8M dated news articles, and in an evaluation of 12 models research tools lower Brier score for all 12, forecasts improve as events unfold with the largest gains at steps where more newly dated evidence is recorded, yet every model still trails historical market forecasts in both Brier score and accuracy, a belief notebook carried between dates lowers research cost but does not consistently improve forecast quality, and the environment additionally provides intera

Source-provided article image: Forecast-Dojo: Replayable Environments for Benchmarking and Training LLM Forecasting Agents
Figure 1 ·

Figure 1: Overview of Forecast-Dojo. Top: resolved Polymarket events are filtered and split by time into training and evaluation events, and CC-News articles are cleaned, dated, and indexed. Middle left: each question is forecast at a fixed sequence of dates as the visible corpus grows, with an optional belief notebook M t M_{t} passed between steps. Bottom left: within one step, the agent searches and reads articles dated on or before τ t \tau_{t} , runs code in a sandbox, and commits a forecast p t p_{t} . Right: the realized outcome Y Y stays hidden from the agent. The evaluator scores each forecast against it, using held-out events for evaluation and training events for learning.

arXiv

Interpretation

Forecast-Dojo combines resolved prediction-market questions with dated news into a replayable environment, letting agents research an event and revisit their predictions at successive historical dates. Unlike conventional forecasting evaluation that must wait for new events to resolve before feedback arrives, pairing resolved questions with historical news lets the same tasks and tools support repeated evaluation, collection of training interactions, and feedback from recorded outcomes. The abstract states the environment's composition: 1,568 Polymarket events split by time into training and evaluation periods, plus 18.8M dated news articles; this is evidence at the level of environment scale and design rather than model internals.

In an evaluation of 12 models, research tools lower Brier score for all 12, and forecasts improve as events unfold, with the largest gains at steps where more newly dated evidence is recorded. This offers a consistent directional result across 12 models on whether retrieval-style research actually helps LLM forecasting, and ties improvement to when evidence arrives rather than reporting only an aggregate score for one model. The evidence comes from the 12-model evaluation and step-level gain comparisons; the abstract does not report per-model Brier values, confidence intervals, or statistical tests, so the strength is directional consistency rather than a precise effect size.

Despite those improvements, every model still trails historical market forecasts in both Brier score and accuracy. This anchors LLM forecasting agents against an external reference point, indicating that gains from research tools have not yet closed the gap with aggregated market forecasts. The abstract states this comparison directly as "Every model still trails historical market forecasts in both Brier score and accuracy," without giving the size of the gap.

A belief notebook carried between dates lowers research cost but does not consistently improve forecast quality, while the environment provides interaction trajectories and outcome feedback for agent learning, with supervised fine-tuning as a proof of concept. The first separates the goals of saving cost and improving quality; the second extends the environment from pure evaluation to training use by supplying reusable interaction trajectories and outcome feedback. The abstract states that the belief notebook "lowers research cost but does not consistently improve forecast quality" and frames supervised fine-tuning as a "proof of concept," that is, an initial demonstration rather than a complete training result.

Perspective

The environment targets LLM agents that must research public events and produce probabilistic judgments, and it supports evaluation, interaction-trajectory collection, and feedback from recorded outcomes, with supervised fine-tuning as a proof of concept; its data consist of 1,568 Polymarket events (split by time into training and evaluation periods) and 18.8M dated news articles, so the results apply to the event types and time range covered by this prediction market and news corpus. For readers who want to reproduce or extend forecasting-agent research, the environment's value lies in reusing the same tasks and tools repeatedly without waiting for new events to resolve.

The abstract does not report per-model Brier scores, the size of the gap versus market forecasts, statistical tests or confidence intervals, nor the specific setup and magnitude of the supervised fine-tuning, so the practical size of the improvement cannot be judged from the abstract. The conditions under which the belief notebook lowers research cost without consistently improving quality, and differences across event types or time periods, would also require the body's breakdowns to confirm. In addition, the text available here is the abstract and bibliographic information; figures and body details are not included, so specific numbers and experimental settings should be checked against the original paper.

Sources