No individual labels needed: reasoning learned from prediction-market prices transfers zero-shot to user simulation
Lead
A language model learns behavioral reasoning from prediction-market price movements, reaching 0.674 directional accuracy on Polymarket and transferring zero-shot to four user-simulation benchmarks without any individual-level annotations.
Story
The training signal for user simulation can come from population-level outcomes instead of individual-level annotations. Earlier simulators either inherited reasoning from pretraining priors or fitted it directly to individual behavioral records, so supervision for behavioral reasoning stayed confined to micro-level data. On SWM-Bench, the model reaches 0.674 directional accuracy and 0.362 correlation on Polymarket, above both the strongest time-series baseline and GPT-5.5.
A forecast is decomposed into readable reasoning steps: the model infers representative groups of market participants, predicts how each interprets the news and updates its beliefs, reasons about their interactions, and aggregates those responses into a price. Training uses GRPO reinforcement learning with a hindsight-regret curriculum: the model is shown the realized outcome and asked only to name the driving groups, and the forecast gain from that hint measures how informative a transition is. Training data are daily prediction-market trajectories from Polymarket and Kalshi in SWM-Bench paired with contemporaneous news, with Qwen3-4B as the backbone.
What to watch
The next step worth taking is scaling the model to larger backbones and other model families, since nearly all experiments share the Qwen3 backbone; a large-scale human evaluation would also test whether the reasoning traces match human judgments. For teams working on user or social simulation, this reasoning can be used directly as a zero-shot simulator or as a synthetic data generator that supplies training examples to a downstream simulator.
The human evaluation used five annotators scoring 20 reasoning traces per model, which is small, and the authors list large-scale human evaluation as future work. The observation that macro forecasting and individual simulation improve together comes from four training checkpoints, and whether it holds on other backbones and larger models still needs checking. Absolute user-simulation scores also depend on the judge model and scoring pipeline, so cross-table comparisons should be made with care.
