PyRUA-Lean lifts GPT-6 Astra robot agents from 63.1% to 71.7% success under equal call budgets, with 49% fewer calls and 65% fewer input tokens on jointly solved tasks
Synopsis
PyRUA-Lean is an interactive code-execution framework in which a VLM agent composes classical robot primitives and frozen VLA policies into Python cells that perform conditional checks and local retries and return only explicitly requested images and state feedback; across 700 simulated task instances from LIBERO-PRO, RoboTwin 2.0, and RoboCasa365, and against a tool-calling baseline using the same GPT-6 Astra planner and the same primitives, it raises overall success from 63.1% to 71.7% under equal LLM-call budgets and uses 49% fewer LLM calls and 65% fewer input tokens on instances solved by both agents.
Interpretation
PyRUA-Lean composes robot primitives into executable Python cells that can branch, retry locally, and compute geometry, with only explicitly printed outputs and requested camera images returning to the VLM context. Code-based robot control (such as Code as Policies and ProgPrompt) and systems with execution feedback already existed, but this work centers on comparing two interfaces, interactive code execution versus tool calling, over the same primitives in terms of success and inference cost. The paper presents the interface design, the primitive list (Table 1), and a per-cell execution model, and its appendix counts behavior over 11,145 cells across 700 episodes: a cell runs 2.1 acting primitives on average and 45% of cells run two or more, whereas one tool-calling LLM call runs 0.65.
Under equal LLM-call budgets, PyRUA-Lean raises overall success from 63.1% to 71.7%, an 8.6 percentage-point or roughly 14% relative gain, with improvements in all four benchmark groups. Gains are 11.0 points on LIBERO-PRO, 9.2 on RoboTwin 2.0, 7.8 on RoboCasa365 atomic tasks, and 5.0 on composite tasks; PyRUA-Lean alone solves 96 instances while tool calling alone solves 36. 700 paired task instances and 1,400 episodes, five environment seeds per task, with both agents using the same GPT-6 Astra planner, the same robot stacks, and the same frozen VLA policies, and neither using cross-episode memory.
On instances solved by both agents, PyRUA-Lean cuts mean LLM calls from 17.0 to 8.7 and input tokens from 788k to 276k, with estimated cost falling from $1.63 to $0.74 per episode. Token decomposition shows the savings come mainly from fewer invocations rather than uniformly smaller prompts: on RoboCasa365 atomic tasks PyRUA-Lean's calls are larger on average, yet fewer invocations still reduce total token usage. Metrics are whole-episode means and ratios of aggregate totals on the jointly solved subset, counting post-success calls and cached input tokens; the authors note this is an accounting analysis that does not independently isolate each interface feature's causal contribution.
Ablations show token efficiency persists without VLA policies or operating guides, but gains vary with task structure and are smallest on VLA-dominated RoboCasa365 atomic tasks. Without VLA policies, PyRUA-Lean reaches 68.8% success on RoboTwin 2.0 versus 42.8% for tool calling, while on RoboTwin 2.0 without VLA policies removing guides raises tool-calling success from 42.8% to 60.4%, indicating guidance effects depend on the interface. Ablations run on every task instance of the main comparison; the authors also note that jointly solved subsets vary across settings, so their token costs do not isolate component effects, and that including failed episodes gives a tool-calling-to-PyRUA-Lean cost ratio of about 1.1 on RoboCasa365 atomic tasks.
Perspective
This work targets simulated robot manipulation settings that use a VLM planner, classical motion and perception primitives, and frozen VLA policies, and it speaks to system designers who want to complete more tasks within the same call budget while controlling which observations enter the model context. Its evidence is at the interface level: conditional checks, local retries, persistent program state, and on-demand image requests inside code cells can reduce model invocations without changing the underlying primitives. The authors state that future work will measure execution latency and recovery on physical robots and examine how primitive granularity and cross-episode reuse of validated routines affect success, inference cost, and adaptability.
Readers should keep in mind that the evaluation is limited to GPT-6 Astra, RPent's robot stacks, and simulation, with one run per agent per task instance, so run-to-run variability, other planners, and physical-robot performance remain open questions. The authors also note that the comparison evaluates the complete interfaces without fully isolating the effects of primitive composition, persistent state, and selective feedback, and that the token decomposition is an accounting analysis. In addition, returning images only on request does not by itself reproduce the gains: on LIBERO-PRO, modifying the tool-calling baseline to request images on demand drops success from 83.0% to 68.5%, suggesting observation requests must be considered together with action execution and replanning. Of the 96 instances solved only by PyRUA-Lean, 52 are cases where the baseline exhausts its call budget and 44 are cases where it ends the episode itself and reports failure, and in about a quarter of those the VLA policy simply succeeded in one run and failed in the other, so the success difference cannot be attributed entirely to better recovery logic.
