AutoGUIWorld synthesized 79,266 GUI trajectories with image generators, lifting OSWorld from 33.0% to 40.8% and ScienceBoard from 14.0% to 32.2% after fine-tuning
Synopsis
AutoGUIWorld introduces a data generation framework that combines the visual priors of image generators with the task knowledge of a planner to synthesize GUI interaction trajectories without deploying or running the corresponding software environments: it samples initial GUI scenes from structured specifications of operating-system context, visual appearance, and interface state, generates tasks conditioned on those scenes, has a planner specify atomic actions and their intended visual consequences, and has an image generator iteratively edit the current screenshot to produce subsequent observations, yielding 79,266 spatially annotated step-level training samples across Ubuntu, Windows, macOS, and Chrome; fine-tuning Qwen3.
Interpretation
The work splits GUI trajectory generation into a planning layer and a visual world-model layer: a Meta Planner produces an ordered sequence of atomic actions and their intended visual changes from the task and fixed seed context, Voyager expands the current screenshot, planned action, and rollout context into a rendering prompt, and Image2 edits the current screenshot accordingly to produce the next visual state, yielding temporally coherent screenshot-action-screenshot trajectories without instantiating real system state. Prior trajectory acquisition relied on human demonstrations, extraction from tutorials and screen recordings, or automated exploration in executable environments, so coverage stayed tied to accessible executable environments or existing records; AutoGUIWorld substitutes pretrained image-generation priors for environment execution, extending coverage from deployed software toward software that can be described structurally. The paper provides a full formal factorization (task and initial visual state sampling, planning layer, visual world-model layer) and a closed-loop rollout description, and releases the prompt templates and action-space definitions for each stage.
The work constructs 79,266 grounded and filtered step-level training samples, comprising 35,209 Chrome, 24,250 Ubuntu, 14,778 Windows, and 5,029 macOS samples, and analyzes their coverage and transition defects. Relative to real corpora such as ScaleCUA, the synthetic data increases detected UI coverage in all twelve domain-threshold comparisons, with relative gains of 63-77% on Ubuntu, 51-75% on Windows, 49-71% on Web, and 10-13% on macOS; meanwhile the synthetic-to-ScaleCUA Qwen embedding MMD ranges from 0.59 to 1.24, on the same empirical scale as variation among independently collected real datasets. Coverage is measured by rasterizing the union of OmniParser-v2.0 detection boxes, and visual alignment by MMD over Qwen3.5-9B merged visual tokens plus a C2ST probe, reported with group-bootstrap confidence intervals; a manual audit of 180 detection overlays found no systematic tendency to label generated texture as UI elements.
In real-environment evaluation, AGW-35B, obtained by fine-tuning Qwen3.5-35B-A3B on AutoGUIWorld trajectories, improves on all four interactive benchmarks, with the equal-weight mean of the four rising from 23.6% to 36.5%, and the largest gains on macOSWorld and ScienceBoard at 16.9 and 18.2 points. This provides direct evidence that synthetic trajectories transfer to real software environments: OSWorld rises from 33.0% to 40.8%, WAA from 19.4% to 27.9%, macOSWorld success from 28.1% to 45.0%, and ScienceBoard success from 14.0% to 32.2%; on ScreenSpot-Pro overall grounding accuracy rises from 31.7% to 57.1%. Evaluation uses fixed task sets and a shared harness, with 361 OSWorld tasks, 154 WAA tasks, 231 macOSWorld tasks, 143 ScienceBoard tasks, and 1,581 ScreenSpot-Pro examples, and reports training curves over six checkpoints; in the comparison against an AgentNet real-demonstration run, AGW-35B exceeds its best evaluated checkpoint on all four benchmarks.
The work quantifies the fidelity boundary of image generation as a world model: across 42,526 evaluable desktop transitions, a VLM audit flags 1,293 action-image inconsistencies (3.04%); after removing steps with an invalid action, absent target, or incorrect grounding, 818 of 39,351 transitions remain inconsistent (2.08%). It separates visual plausibility from faithful realization of an action and reports the residual error distribution by action type: 7.58% for drag, 7.09% for scroll, 6.35% for text entry, 3.23% for hotkeys, 2.16% for key presses, and 0.67% for clicks, identifying exact content preservation as a distinct challenge for visual transition synthesis. The audit uses a structured five-dimension check (target existence, box hit, box tightness, action validity, observation transition consistency), re-grounds 3,390 target boxes before rechecking, and reports per-OS results plus representative failure modes.
Perspective
This work targets researchers and engineering teams that need large-scale GUI interaction trajectories to train agents, and it applies to expanding training-data coverage without deploying or running the corresponding software environments, especially for specialized software and workflows that are hard to install and configure. It enables follow-up work to reuse the pipeline of structured sampling space, seed-conditioned task generation, planner-guided rollout, action grounding, and quality filtering across four desktop and browser domains, and to convert generated trajectories into supervised fine-tuning, trajectory replay, or reinforcement-learning data. The paper also notes that domain-specific training could yield vertical GUI world models tailored to particular software or workflows, improving long-horizon consistency and supporting further data generation within those domains.
Image generators may accumulate errors over longer trajectories and more involved interactions, producing hallucinated interface states or action outcomes; the reported residual inconsistencies concentrate in text entry, dragging, scrolling, and keyboard interactions, including command substitution, invented document content, premature formatting, missing action effects, and unintended persistent-content changes. In the current release, precondition and blocker fields are zero throughout and rendered terminal states are almost always successful, so coverage of failure paths remains to be seen. Trajectory structure statistics (such as 11.9 and 11.4 average steps on Ubuntu and Windows and 0.92 action-type entropy on macOS) describe the synthetic corpus, and matched real-trajectory metadata are unavailable for that analysis. The comparison with the AgentNet run uses different inference settings, and the training-step axis does not normalize epochs, tokens, or compute across runs. In addition, the loaded text is the full paper text, and some figures and appendix visualizations appear as textual descriptions, so checking specific graphical details still requires the original.
