OpenRUA turns an off-the-shelf coding agent into a zero-shot visuomotor policy via ROS 2 terminal access alone, reaching 99.0% on CaP-Bench and 87.0% on LIBERO-PRO
Related research and updatesSynopsis
OpenRUA introduces a zero-abstraction harness that gives an off-the-shelf coding agent (Claude Code powered by Claude Opus 5) only terminal access to the robot's native software interface ROS 2, plus ROS 2 documentation and basic tools, with no prescribed workflow or bespoke interfaces, recasting perception as file I/O and manipulation as coding; it achieves 99.0% success on CaP-Bench and 87.0% on LIBERO-PRO, showing that an off-the-shelf coding agent can serve as a zero-shot visuomotor policy without bespoke primitives or task-specific training.
Figure 1: OpenRUA: a minimalist workspace-as-harness for robot-use agents. Given a task such as placing the block in the bowl, OpenRUA enables any off-the-shelf coding agent to operate a real or simulated robot through its native ROS 2 interface, bypassing bespoke abstraction layers. The workspace serves as the harness by only providing ROS 2 documentation and basic tools, making robot perception accessible through file I/O and manipulation through coding, without orchestrating any agentic workflow. Within this workspace, the agent exhibits emergent perception and control capabilities as it works to complete the task. It writes and executes ros2cli commands or rclpy programs that send control requests to the robot’s controllers, here grasping and lifting the wooden block. The agent saves new RGB-D observations as workspace files and uses images, computed measurements, and terminal feedback to autonomously verify progress and refine its actions until the block is placed in the bowl.
arXivInterpretation
OpenRUA is a zero-abstraction harness that bypasses bespoke abstraction layers by giving an off-the-shelf coding agent only terminal access to the robot's native software interface ROS 2. Prior work primarily engineered complex custom harnesses to orchestrate agents for robot use, prescribing specialized workflows and providing bespoke interfaces; OpenRUA replaces this with a minimalist workspace-as-harness design that offers only ROS 2 documentation and basic tools, orchestrating no agentic workflow and leaving the coding agent to organize its own work. The abstract states this contrast with the prior custom-harness route and describes the design as offering only ROS 2 documentation and basic tools; no quantitative comparison against a baseline harness is given.
Under this minimalist design, OpenRUA with Claude Code powered by Claude Opus 5 achieves success rates of 99.0% on CaP-Bench and 87.0% on LIBERO-PRO. This indicates that an off-the-shelf coding agent can act as a zero-shot visuomotor policy through the robot's native interface without bespoke primitives or task-specific training, addressing the question of whether additional harness engineering is necessary. The abstract reports success-rate numbers on two named benchmarks; sample sizes, repetition counts, and statistical uncertainty are not provided.
Analysis reveals emergent behaviors of the coding agent: for perception, it spontaneously writes programs that process raw sensory inputs and derive metric measurements in 96.80% of episodes; for manipulation, it spontaneously builds motion-control clients (e.g., gripper control) in 95.87% of episodes and closed-loop control programs (e.g., adjusting motion based on sensor feedback) in 50.13% of episodes. These proportions quantify how far the agent spontaneously forms perception and manipulation programs without bespoke primitives or workflow orchestration, adding mechanism-level observations beyond the success rates. The abstract gives three per-episode percentages; the total number of episodes, the counting criteria, and confidence intervals are not stated.
Perspective
The work targets researchers and engineers who want coding agents to operate robots directly without building custom abstraction layers; the applicable setting is a robot exposing its native ROS 2 interface together with an available off-the-shelf coding agent (in the abstract, Claude Code powered by Claude Opus 5). In that setting, OpenRUA offers a path whose entire provision is ROS 2 documentation and basic tools, with the agent organizing its own work, and its reported success rates and emergent-behavior proportions can serve as a starting point for reusing this minimalist design on comparable native interfaces.
The abstract does not state the task composition of CaP-Bench and LIBERO-PRO, the total number of episodes, the evaluation protocol, or repetition counts, nor does it give direct comparison numbers against a custom-harness baseline, so the comparable range of 99.0% and 87.0% still needs confirmation in the full text. The counting criteria behind the three emergent-behavior proportions (96.80%, 95.87%, 50.13%) are likewise not expanded in the abstract. In addition, the abstract reports results only for Claude Code powered by Claude Opus 5; whether other coding agents or models show the same behaviors and success rates remains an open question. Because this reading is based on the abstract only, figures and experimental details are not included, and the full context of these numbers should be checked against the original.
