Skip to main content
Back to timeline
arXivSource publication:

RecreationWorld: Stitching Interface Operation and Code Writing into One Long-Horizon Trajectory via Recreation

Synopsis

This work introduces RecreationWorld, a framework, and RecreationBench, a benchmark: given a running reference application, an agent must discover its behavior and deliver a buildable, runnable recreation, with reproducible environments on Ubuntu, macOS, Windows, Android, and Web plus a unified GUI-and-coding harness, automatic scoring by reference-derived hidden programmatic and visual assertions, and 35,000 filtered recreation trajectories for training; evaluation shows GPT-6 Astra leading at 58.1% overall while passing all programmatic tests on just 2.8% of tasks, and recreations reproducing static interface structure more reliably than interactions and computed outputs.

AI-generated editorial illustration: RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents

Interpretation

It defines a hybrid computer-use task built around recreation: the specification is encountered only by operating a running reference, while the deliverable is source code that must build and run, forcing an explore–implement–verify loop. Prior benchmarks tend to emphasize one interaction mode (GUI-only or terminal-only); this work places both modalities inside a single long-horizon task without prescribing action order or modality. The task definition and five-platform delivery contracts (desktop source with build/launch, Android Gradle producing an APK, Web self-contained index.html) appear in the main text and Table 1; trajectory statistics report a median of 282.5 top-level calls and 9.08 GUI–code-edit transitions per 100 calls on RecreationBench.

It builds scalable, verifiable environments and an evaluation layer: the running reference acts as an oracle from which hidden behavioral tests are derived, with programmatic assertions reading AT-SPI, AXUIElement, UI Automation, and UiAutomator, and visual assertions judged by a fixed VLM at frozen checkpoints. It operationalizes the asymmetry of verification across five platforms in a reproducible pipeline, scoring independently of the candidate's language, framework, or architecture. Every proposed assertion must first pass on the reference and undergo human review before the suite is frozen; the coverage audit reports about 77% of cases leaving the start surface, 24% traversing two or more surfaces, 94.2% checking an interaction outcome, and 40.7% requiring an exact expected result; 81 tasks package 393 fixture files.

Training on recreation trajectories transfers beyond recreation: 7,000 selected trajectories per platform form a balanced 35,000-trajectory SFT mixture used to fine-tune two model initializations, both finishing above their first evaluated checkpoints on five out-of-distribution benchmarks with gains of up to 17.9 percentage points. It offers initial evidence that recreation supervision improves broader coding and hybrid computer-use capabilities rather than only recreation itself. Five OOD benchmarks (ProgramBench, GameCraft-Bench, Vision2Web, OSWorld 2.0, WeaveBench) with one trial per task at each checkpoint; the sweeps are not uniformly monotonic, and behavioral metrics show increases in own-output image reads and GUI observation and interaction, which the authors describe as descriptive rather than causal.

It profiles current frontier models and their failure modes: static interface structure is reproduced more reliably than interactions and computed outputs, recreations are substantially smaller and more structurally concentrated than references, and the final source edit often reaches submission without a relaunch and inspection. It decomposes the score gap into locatable category gaps and process metrics rather than a single aggregate number. Across 16 model–platform comparisons on Ubuntu, macOS, Windows, and Android, structure has the highest mean pass rate in 15, with the weakest category trailing by 9.1–34.4 points; Web static content exceeds interaction by 28.2–36.7 points; 89.4% of recreations contain less production source than their references with a median LOC ratio of 16.9%; the strict final edit–relaunch–inspection sequence appears in only 23.6–47.5% of trajectories.

Perspective

The framework targets pinned, reproducible offline reference applications: desktop and Android use a source-blind setting, while Web protects ground truth and tests instead because a browser necessarily receives the client code; the task pool comes from open-source applications and benchmark-authored sites across Ubuntu, macOS, Windows, Android, and Web. It suits studying the long-horizon explore–implement–verify loop, generating automatically scored training experience, and comparing how models coordinate interfaces and code. Workflows involving live services, changing external state, or real-user interaction are outside the current setting.

The hidden suites sample a finite set of states and interaction paths, so passing them does not establish equivalence across all states; public reference implementations may have appeared in pretraining data, and line-level overlap screening cannot rule out renamed, refactored, or fragment-level memorization; Android does not attest a packet-level egress filter, and an installed reference package remains recoverable in principle; Web's scrape-and-replay detection has a boundary, since reference content preserved through an intermediate schema can escape it; the programmable-runtime comparison is a one-rollout end-to-end configuration comparison and cannot isolate the causal effect of the persistent runtime; and although this reading covered the full text, some numeric details in figures and appendices are summarized here in prose, so exact reproduction should consult the original figures.

Sources