Skip to main content
Back to timeline
arXivSource publication:

A2Z GameSpec-Bench tests coding agents on 100 long-form GDDs: near-99% verifiable rate, but top overall GDD Fidelity only 77.0

Synopsis

The work introduces A2Z GameSpec-Bench, a benchmark of 100 long-form game design documents (50 Small and 50 Big, averaging 14,085 and 26,297 tokens) that turns each GDD into a fixed Dependency-Aware Contract and evaluates coding agents' end-to-end game development faithfulness through source-code inspection, scenario-based replay, and adaptive playtesting, finding that compilable and runnable outputs do not imply faithful implementation (e.g., Claude-Fable-5.1 reaches a 98.7% verifiable rate but an overall GDD Fidelity of 77.0), that dependency context raises coverage of recorded playtest violations from 71.1% to 80.2%, and that requirement-specific feedback improves overall GDD Fidelity by 10.9% relative to self-revision after two rounds on Big GDDs.

AI-generated editorial illustration: A2Z GameSpec-Bench: How Faithfully Can Coding Agents Generate Games from Game Design Specifications?

Interpretation

The benchmark converts each long-form GDD into a fixed Dependency-Aware Contract in which rules record conditions, triggering events, and expected effects, invariants record constraints within a stated scope, and directed edges connect a rule that updates a state referenced by another rule's preconditions or emits an event that triggers it. Existing game-development benchmarks typically use compact specifications (185 tokens for GameDevBench, 131 for PlaytestArena, 1,547 for GameCraft-Bench in the comparison table) and do not explicitly represent dependencies among requirements; this work uses long-form GDDs averaging 20,191 tokens and builds an explicit graph. The contract is constructed and frozen before any build is inspected and stays fixed across agents and revision rounds; across 100 GDDs, repeated contract construction reaches 93% citation overlap on Big (57% for a naive judge), and three rule generations preserve reachability agreement, dependency-edge Jaccard similarity, and role preservation each at or above 0.999.

Evaluation binds judgments and evidence from three channels—source-code inspection, scenario-based replay, and adaptive playtesting—to the same requirements, with replay using a fixed input policy plus adaptive frame selection and playtesting using Code-as-Policy bots for normal and adversarial tests. Existing benchmarks usually cover only one or two of these channels (only A2Z GameSpec-Bench is marked across source code, rendered behavior, and adaptive play in the comparison table), and replay is often scored on fixed-interval frame sampling. On identical recorded replays, visual rubrics, and frame counts across 50 Small games, adaptive frame selection is preferred in 60.5% of comparisons and wins 35 of 50 games (13 for fixed-rate sampling, 2 ties); on 86 dependency-linked rules, a second normal pass raises judgment coverage from 40.5% to 41.0%, while an adversarial pass raises it to 48.5%.

Compilability and runnability diverge from faithful implementation, and locally passing rules can depend on failed prerequisites. It separates 'it runs' from 'it runs as designed' and provides a graph-based reachability proxy for the latter. Claude-Fable-5.1 achieves a 98.7% verifiable rate under compile and runtime checks while its overall GDD Fidelity is 77.0; across 100 GPT-5.6-Sol games, the mean pass rate for dependency-linked rules is 72.4% from local source-code judgments but 22.7% when all upstream rules must also pass; among rules with a full source-code score only 76.8% are confirmed satisfied during playtesting, while 25.7% of rules with zero source-code score are judged satisfied during playtesting.

Dependency context directs inspection toward runtime failures that enumerated evaluation misses, and requirement-level feedback yields a measurable revision gain. Relative to an enumerated baseline that scores GDD-derived requirements individually, adding dependency edges expands both inspection targets and violation coverage; relative to self-revision that only reads the GDD and source code, three-axis feedback improves more. On one initial GPT-5.6-Sol build per each of 100 GDDs, dependency context raises violation coverage from 398/560 (71.1%) to 449/560 (80.2%), a 9.1-point gain with a 95% paired game bootstrap interval of 6.1–13.0, adding 51 violations across 26 games; on 50 Big GDDs, Source + Replay + Playtest reaches a mean overall GDD Fidelity of 74.3 after two rounds versus 67.0 for self-revision, a 7.3-point (10.9% relative) gain.

Perspective

The benchmark targets end-to-end generation of 2D single-player browser games (Phaser 4.1, TypeScript, Vite) from 100 synthetic GDDs split into 50 Small (averaging 14,085 tokens and about 54 outcome requirements) and 50 Big (averaging 26,297 tokens and about 84 outcome requirements); each GDD specifies on average 69 outcome requirements and 32 invariants, and evaluation uses 20 replay scenarios. It is suited to comparing coding agents' faithfulness under long-form specifications, pinpointing requirement-level deviations, and serving as a source of revision feedback; because the framework separates GDD-derived contracts from engine-specific interfaces, it can extend to other runtimes, demonstrated with three Three.js 3D games (Sonic: Cascade Coast, Diablo Cathedral, Rocket League).

Contract construction and evidence interpretation rely on generative models, so fixed contracts and repeated judgments cannot fully eliminate omissions or evaluator bias; main results are limited to synthetic GDDs and the 2D Phaser setting, and the 3D demonstration covers only three games, which is not enough to establish behavior across engines and genres. Dependency edges are defined as potential prerequisites, so the reachability proxy indicates possible blocking rather than demonstrated runtime unreachability. In addition, although this evidence bundle is full text, some figures and appendix details appear in summarized form, so checking specific numbers and per-requirement verdicts would still require the original appendices.

Sources