Kaggle Game Arena pits ten frontier models against each other in chess, poker and Werewolf: Gemini 3 leads chess and Werewolf, GPT-5.2 tops poker at +46.6 BB/100, and rankings do not agree across games
Synopsis
This technical report introduces Kaggle Game Arena, an open head-to-head evaluation platform in which ten frontier models play large-scale round-robin matches in Chess, Poker and Werewolf through a uniform text harness (900,000 poker hands, 180,000 per model), ranking them by objective outcomes such as wins, chip counts and role-level decomposition rather than subjective judgment, and reports sharp within-game stratification alongside rankings that do not agree across games.
Interpretation
The platform moves evaluation from static question sets to dynamic play: models face off in structured environments whose gameplay strength rises as models improve, avoiding saturation, and rankings rest on objective outcomes such as wins and chip counts rather than human or model judges. Relative to fixed test sets such as MMLU, GSM8K and HellaSwag, and to preference-based evaluation such as Chatbot Arena and MT-bench, this work anchors on ground-truth game outcomes to counter saturation, data contamination and judge subjectivity. The paper documents a uniform text harness, retry handling for invalid actions, variance-reduction techniques and bootstrapped confidence intervals, and releases the harness and full gameplay trajectory dataset.
Chess shows clear stratification: Gemini 3 Pro Preview (internal Elo 1325) and Gemini 3 Flash Preview (1297) lead, followed by o3 (1009) and GPT-5.2 (933), with the Claude 4.5 series at the rear (Opus 236, Sonnet 189, Haiku 122). Using Stockfish centipawn-loss paths converted to win probabilities, the paper localizes the gap to the middlegame and endgame rather than the opening, and finds weaker models increasingly make illegal moves (rethink events) as games progress. Each model pair plays 40 games with color balance (20 white, 20 black), and a Chess Opening variant starting from 20 popular Lichess two-ply openings serves as a robustness check whose rankings match Chess Text.
In poker, GPT-5.2 leads at +46.6 BB/100, ahead of o3 (+29.7) and Grok 4 (+27.1), while GPT-5 mini (-94.9) is a clear outlier; preflop styles vary widely, and preflop aggression does not map linearly onto profitability. The paper introduces duplicate poker (hand mirroring) to LLM poker benchmarking and lets models explicitly model opponents over complete text histories within 100-hand episodes, with both players' hole cards revealed after each hand. Each matchup spans 20,000 hands and the ten-model round robin totals 900,000 hands (180,000 per model), with block-bootstrap confidence intervals; GPT-5.2's 95% interval does not overlap those of o3 or the middle- and bottom-tier models.
In Werewolf, Gemini 3 Pro Preview and Gemini 3 Flash Preview achieve the highest net ratings with balanced positive contributions across roles, while GPT-5 mini shows severe negative contributions in every role; substitution analysis finds Gemini 3 Pro Preview is always an improvement regardless of role or model replaced. The paper applies game-theoretic evaluation (GTE) to decompose overall skill into role-specific contributions (Werewolf, Seer, Doctor, Villager), avoiding the conflation of individual skill with structural base win rates in traditional Elo or OpenSkill ratings. Rule ablations via 500 self-play games per configuration suppress the Villager win rate from 73.4% under standard rules to 56.7%, with bootstrapped confidence intervals and an acyclic tournament graph confirming transitive dominance.
Perspective
The work targets evaluators and model developers who need continuous, reproducible comparison of frontier models' strategic capability. It applies to three structured settings: Chess (perfect information), Poker (imperfect information, two-player zero-sum) and Werewolf (multiplayer, general-sum, information-asymmetric), with the harness and full gameplay trajectory dataset released for downstream reuse. Its design intent is that new games, variants and evaluation methodologies can be added over time without invalidating prior results, supporting longitudinal measurement and research on benchmark design itself.
Open directions the paper itself names include: a fixed number of games per model pair rather than uncertainty-driven adaptive compute allocation; model churn that complicates longitudinal comparison and fragments the match graph; game-specific primary metrics with no consolidated cross-game meta-rating yet; and open questions about fair credit assignment and aggregating team outcomes in multiplayer games. In addition, Werewolf balance uses 500 self-play games by a single high-capability model as an empirical proxy rather than an exact Nash equilibrium; the Stockfish external calibration for chess is less reliable outside the engine calibration range; and Werewolf heuristic metrics (KSR, IRP, VSS and others) lack discriminative variance, which is why GTE is treated as definitive. Readers citing specific rankings should read them together with the reported confidence intervals and these scope conditions.
