RSIGame pairs local explore-diagnose-improve with a global best-checkpoint loop, lifting Qwen3.8-27B from 37.07 to 61.38 on Godot and past GPT-5.5's 50.26 one-shot score
Synopsis
RSIGame organizes automatic game development as recursive self-improvement: a local explore-diagnose-improve loop broadly explores the executable game, diagnoses and prioritizes issues, and performs evidence-grounded revision while an evolving checklist accumulates testing and improvement guidance; a global loop tracks overall quality, preserves the best checkpoint, and detects saturation or regression; and successful development experience is internalized into the generator through supervised fine-tuning, consistently improving game quality across 140 GameCraft-Bench tasks, two engines, and five generators under matched development budgets, with experience internalization enabling Qwen3.8-27B to reach 61.38 on Godot and 58.53 on Phaser, exceeding GPT-5.
Interpretation
RSIGame splits game development into complementary local and global loops: the local loop has a controller choosing the exploration direction, an explorer interacting with the executable game to collect behavioral evidence, an editor proposing edits, and a verifier replaying the build to check whether the targeted issue is resolved without regressions, writing outcomes back into an evolving development checklist; the global loop's game quality monitor compares the retained champion against each candidate at checkpoints, keeps the better build, and declares saturation once the champion survives three consecutive checkpoints. Prior systems such as Play2Code, VibeGame, and OpenGame already include iteration or verification roles, but the paper argues they treat iteration itself as the source of improvement; RSIGame's difference is organizing iteration into a controlled process with direction, prioritization, best-checkpoint retention, and saturation-aware stopping, and explicitly isolating the monitor from the benchmark evaluator by reading only a separate proxy rubric. Across 140 GameCraft-Bench tasks, two engines, and five generators, RSIGame improves over the frozen initial project by roughly 8.77 to 14.26 Overall points and over Play2Code by roughly 1.55 to 13.79 points under matched development budgets (up to 26 tool calls per round, 30 rounds); on Godot with GPT-5.5, RSIGame versus Base is +14.26 (95% CI [11.49, 17.06]) and versus Play2Code is +13.79 (95% CI [10.89, 16.65]).
The paper reports that global quality monitoring makes development-time scaling more reliable: simply extending the budget does not guarantee better games, as Play2Code quickly plateaus and often regresses, whereas RSIGame converts additional development rounds into higher game quality across both engines and both strong and weak initializations; the retained checkpoint closely tracks the oracle best within each budget, and saturation-aware stopping reaches comparable quality to much longer fixed-budget runs with fewer rounds. This separates 'iteration can improve' from 'iteration can keep improving': the paper uses Figure 3 and Figure 4(a)(b) to show that the local loop alone produces volatile trajectories, while the global monitor stabilizes the process by retaining the best state and makes the stopping rule part of compute efficiency. The paper reports that the retained checkpoint closely tracks the oracle best within each budget and that saturation-aware stopping achieves comparable quality to longer fixed-budget runs with fewer rounds; checkpoints are taken every 3 rounds and development stops once the champion is unchanged for 3 consecutive checkpoints.
The paper internalizes development experience into the generator: using GPT-5.5 generation traces, distilled planning traces, and independently verified improvement rounds from GLM-5.3-Flash, it applies LoRA supervised fine-tuning to Qwen3.8-27B so the model learns not only final games but the intermediate planning, tool-use, diagnosis, and revision decisions. Relative to test-time-only improvement, the paper extends self-improvement from context optimization to parameter optimization; the ablation shows training on generation traces alone yields uneven gains, while adding planning and verified improvement experience produces consistent improvements across Mechanics, Depth, Visuals, and Art, especially in Mechanics and Depth, with roughly an 11-point Overall gain before any test-time development. The training corpus contains 2,213 generation trajectories, 2,108 planning traces, and 2,013 verified improvement rounds drawn from 4,003 candidates; after fine-tuning, Qwen3.8-27B's one-shot Godot score rises from 37.07 to 48.22, and with RSIGame reaches 61.38, exceeding Codex GPT-5.5's 50.26 one-shot score while generation tokens fall from 6.41M to 0.57M.
The paper audits local improvement reliability with independent adjudication: evidence-grounded pre-improvement verification raises grounded precision from 58.6% to 72.3% and cuts ungrounded targets per round from 1.93 to 0.50; post-improvement replay verification detects 76.2% of unsuccessful improvements over 48 adjudicated rounds with 84.4% balanced accuracy, whereas a build-only baseline detects 0.0%. The paper separates verification into pre-edit and post-edit stages and shows that checking only whether a build runs cannot catch behavioral failures, giving an auditable criterion for whether a round's edit earned its commit. The audit uses Claude Opus 5 to independently adjudicate development trajectories from 40 GPT-5.5-initialized games, with each case evaluated twice and disagreements manually reviewed; in free-play evaluation RSIGame wins 17 of 20 tasks against Base and 18 of 20 against Play2Code, agreeing with the benchmark ordering in 45 of 59 decided comparisons.
Perspective
This work targets agentic development pipelines that turn natural-language specifications into executable games, applies to relatively compact games developable within practical agent budgets, and is validated on two engines, Godot and Phaser. It lets development teams and researchers convert more rounds into higher quality under matched budgets and lets smaller models take on stronger generation tasks through experience internalization; the paper also offers an optional high-level director interface where a human or stronger model supplies direction after autonomous improvement saturates, and reports that later stages are consistently preferred in blind pairwise play during multi-stage development.
The paper states its evaluation centers on Mechanics, Depth, Visuals, and Art, and does not fully capture originality, narrative quality, long-term player engagement, or subjective enjoyment; the high-level guidance mechanism is evaluated on a limited set of multi-stage cases, leaving open when guidance should be invoked, how much guidance is beneficial, and how human and model guidance differ; extending to substantially larger projects, longer horizons, and more complex cross-system dependencies remains an open direction. In addition, replay and scoring vary by task family: across 529 frozen artifacts scored three times, the median range is 0.73 points, with sports, horror, and openworld varying most while racing, rhythm, and strategy have a median range of exactly zero, so single comparisons on high-variance families should be read with care.
