Skip to main content
Back to timeline
arXivSource publication:

HybridCUA lets a 9B agent reach 53.6% on OSWorld by mixing GUI and CLI, 14.8 points above its base model

Synopsis

The work builds HybridCUA-8K, a data pipeline of 5,000 GUI-only, CLI-only, and interleaved GUI–CLI trajectories plus 3,000 verified RLVR tasks, and a two-stage framework of supervised fine-tuning followed by online reinforcement learning with CLI-aware rewards; the resulting HybridCUA-9B reaches 53.6% accuracy on OSWorld (14.8 points above its Qwen3.5-9B base, with average steps falling from 31.6 to 14.0) and 36.0% on WindowsAgentArena (a 4.0-point gain).

AI-generated editorial illustration: HybridCUA: Learning to Orchestrate GUI and CLI for Computer-Use Agents

Interpretation

The paper shows that simply handing a shell to existing agents hurts: once the CLI is exposed, four representative agents lose 2.5 to 11.5 percentage points on OSWorld, and Qwen3.5-27B and EvoCUA-32B route only 15.0% and 0.15% of their steps through the CLI. Prior work either used GUI only or relied on application-specific APIs and tools; this work treats knowing when and how to use the CLI as a capability to be learned rather than assumed. The claim rests on measured comparisons of four existing agents on OSWorld plus a direct CLI step-share metric, making it diagnostic evidence that the bottleneck is interface selection rather than interface availability.

The authors build HybridCUA-8K: 5,023 supervised trajectories (GUI-only converted from UI-MOPD into PyAutoGUI calls, CLI-only sampled from Qwen3.8-27B under a Claude Code harness, interleaved from free-form switching and GUI-to-CLI rewriting with only successful replays kept) plus 3,000 RLVR tasks with executable verifiers, each labeled by ranking 16 rollouts under each of the three interface modes. GUI, shell, and code corpora were previously siloed, with CUA datasets containing almost exclusively GUI actions; this pipeline grounds interface selection and command execution in GUI context and makes whether CLI use offers an advantage an explicit label. Dataset size, sources, trajectory length (10.4 logical steps on average, median 7, 90th percentile 22, maximum 50), and the labeling procedure (16 rollouts, ties within 5 percentage points broken by median successful-rollout step count) are documented in the main text and appendices.

Two-stage training makes interface choice an explicit signal: SFT on the 5,023 trajectories with next-token prediction, then GRPO-based RLVR in a live dual-interface environment, with a trajectory-level reward R_CLI = I[Success]·I[b(τ)=b*] teaching when to use the CLI and a step-level reward r_exec assigning −1 to shell execution failures teaching how to use it. Earlier hybrid-interface work either scored with a coarse VLM or supervised task outcomes alone, leaving interface selection implicit in imitation; here routing and execution reliability become separate reward terms. Ablations show that removing R_CLI costs little accuracy but stalls the efficiency gain (CLI usage 58.9% instead of 64.0%, trajectories shortened by 18.2% instead of 29.3%), while removing r_exec leaves accuracy and steps nearly intact but lets execution errors climb back to 16.5%, versus 11.5% for the full model.

HybridCUA-9B reaches 53.6% accuracy with 14.0 average steps on OSWorld, beating comparably sized models and exceeding AutoGLM-OS-9B by 4.7 points, ToolCUA-8B by 6.8 points, and the four-times-larger UltraCUA-32B by 9.9 points; it reaches 47.1% on OSWorld-MCP and 36.0% on WindowsAgentArena. GUI–API routes require per-application construction, whereas the shell needs none; the model also issues PowerShell commands on Windows despite training only on Linux shells. Main results cover all 361 OSWorld tasks with a 50-step budget and report mean evaluator scores; transfer results are reported separately on the 361 OSWorld-MCP tasks and the 154 WindowsAgentArena tasks and are not combined into a single aggregate.

Perspective

The result targets desktop applications and operating systems that expose usable and stable command-line interfaces, with training data concentrated on the applications included in OSWorld (11 domains, dominated by the three LibreOffice applications and multi-application splits); evaluation runs in sandboxed virtual machines that are reset between episodes and have no access to real accounts or credentials. For a reader, it offers a hybrid-interface route that does not depend on per-application APIs: the GUI handles visual interaction and spatial adjustment, while the shell handles file and system-level operations, and the approach carries over to an MCP-enabled environment and to Windows. The authors also state they will release the trajectories, RLVR tasks, data generation and training pipelines, and the HybridCUA-9B model, making it possible to reproduce and extend on other applications.

Open questions remain: the benefits depend on the availability and stability of command-line interfaces, so applications without usable CLIs, restricted-shell environments, or operating-system-specific command semantics may need substantially different routing behavior; data construction currently focuses on applications covered by OSWorld, limiting representativeness for specialized applications and broader out-of-distribution environments; OSWorld-MCP shares the same task set as OSWorld, so it measures an environment change rather than unseen tasks; reported numbers are mean evaluator scores, and some verifiers return fractional scores, so they are not directly comparable to strict all-or-nothing success rates; each number comes from a single evaluation run and no validation set was used for model selection. In addition, shell access amplifies what a single action can do, and real-world deployment behavior and safety are beyond the scope of this study, requiring further evaluation, safeguards, and human oversight.

Sources