CUA-SWE tested 105 executable tasks and found that giving agents the running interface lifted GPT-6-Astra's four-domain mean from 11.3% to 59.9%
Synopsis
The authors introduce CUA-SWE, a benchmark, interactive development environment, and evaluation pipeline spanning 105 tasks across Web, Game, DevOps, and Mobile, in which an agent edits code, runs commands, operates the running software, and inspects screenshots within a single task, with task-specific deterministic tests deciding the outcome; comparing code-only and hybrid CUA conditions across frontier models, GPT-6-Astra reaches the highest four-domain mean of 59.9%, and every frontier model in the four-domain comparison achieves higher aggregate success with hybrid access, with gains concentrated on tasks whose requirements must be recovered from application materials such as drawings, service contracts, or graphical reference cards.
Interpretation
CUA-SWE places code editing and software use inside one development episode: the agent can read and modify source, run builds and scripts, and operate the running application through mouse, keyboard, and screenshots, while a deterministic verifier independent of the trajectory checks requested functionality, preserved existing behavior, and restrictions on permitted changes. Prior coding benchmarks such as the SWE-bench line and computer-use benchmarks such as OSWorld and WebArena largely assess code changes or interface operation separately, and visual software benchmarks concentrate on particular development domains; this work unifies source-level execution with graphical interaction in the same task and supplies executable correctness criteria. The paper gives a problem formulation as a finite-horizon POMDP, an action-space definition, and a verifier construction process validated in both directions with gold repairs and plausible negative repairs; tasks are LLM-assisted authored, then human reviewed and executably validated.
The benchmark contains 105 tasks: 36 Web, 29 Game, 20 DevOps, and 20 Mobile, annotated by specification provenance as source-specified (S) or application-material-dependent (M), with 63 M tasks (58 using runtime-supplied materials and five using bundled graphical reference cards). Treating where the requirement information lives as an annotatable task property lets the benchmark separate implementation and debugging from recovering a specification through the running interface, rather than assigning a scalar difficulty label. The paper lists tasks per domain with per-task annotation basis and explains why tasks such as 2048, Fabric, and Flappy Bird remain S despite containing runtime fixtures.
In the single-attempt four-domain comparison, GPT-6-Astra reaches the highest mean of 59.9%, 17.7 percentage points above GPT-5.6-Sol, and achieves the highest observed hybrid success in every domain (Web 66.7%, Game 37.9%, DevOps 80.0%, Mobile 55.0%); every frontier model in the four-domain comparison has higher aggregate success with hybrid CUA, with GPT-6-Astra gaining 48.6 percentage points. It turns the question of whether visual feedback actually helps software engineering into a paired measurement stratified by domain and information requirement, and shows the gains are not uniform. The same tasks and the same protected tests are paired across both conditions; the paper reports 32 model-domain cohorts and 840 paired model-task observations, plus 20,000 paired bootstrap resamples on Web and Game.
Gains concentrate on M tasks: across eight frontier models the mean hybrid advantage on M tasks is 42.0 points on Web, 23.6 on Game, 48.1 on DevOps, and 38.8 on Mobile, while on S tasks the mean differences are 0.0 points on Web and on Game; trajectory analysis associates GUI re-verification of the changed application after the final edit with higher success (for example 45.8% versus 5.7% on Web and 83.7% versus 10.6% on Mobile). It splits what interface access contributes into two roles: exposing materials that determine the implementation, and letting the agent exercise the changed application to inspect the consequences of its edits. The stratified comparison uses exactly the same task sets; the trajectory analysis draws on interaction counts from 2,040 evaluation attempts and detailed annotation of 214 trajectories, and the authors describe these as descriptive associations rather than significance tests.
Perspective
The benchmark targets locally hosted applications and games that can be started reproducibly, with tasks confined to Web, Game, DevOps, and Mobile, and evaluation conducted under fixed action budgets and time limits; the conclusions therefore apply to the question of whether an agent, given a particular interface and budget, can turn visual evidence into a software change that passes verification. It speaks directly to two audiences: researchers and engineering teams judging the usability of hybrid CUA agents in real development workflows, and post-training researchers who want to build training signals from executable verifiers. The paper also offers a reusable supervision path: constructing action-prefix data from verified reference replays and applying SFT to Qwen3.8-27B raises the native-verifier pass rate from 11.1% to 33.3%, and it describes a GRPO-based RLVR objective with a parallel rollout interface.
Worth watching: the paper describes the SFT result as early positive indications, and whether the environment provides a verifiable learning signal and scalable RLVR training remains future work, involving parallel rollout collection, stable optimization over long multimodal trajectories, and credit assignment across code edits, observations, and interactions. On trajectory annotation, the authors report low agreement for some stage labels such as incorrect interpretation and mistaken diagnosis, so those labels are treated as exploratory; the association between final-edit GUI re-verification and success is likewise a descriptive estimate in which task difficulty, model, and the trajectory-validation component may also contribute. Repeated attempts show coverage rising with more tries (for example GPT-6-Astra on Game: pass@1 37.9%, pass@3 62.1%) while only 34.5% of tasks succeed in all three attempts, so single-attempt results and stable competence should be read separately; Core Ball and the Mobile schedule family remain unsolved in the single-attempt evaluation, indicating that coupled state and dense dependencies are open difficulties. This summary is based on the paper text and its appendices; the experiments were not independently reproduced.
