Skip to main content
Back to timeline
arXivSource publication:

PyRUA-Lean lifts GPT-6 Astra robot-agent success from 63.1% to 71.7% while cutting input tokens by 65%

Related research and updates

Synopsis

The authors introduce PyRUA-Lean, an interactive code-execution framework that couples feedback-driven primitive composition with selective observation, letting the agent compose classical robot primitives and learned vision-language-action (VLA) policies into Python cells that perform conditional checks and local retries and return only explicitly requested images and state feedback for replanning; across 700 simulated task instances from LIBERO-PRO, RoboTwin 2.0, and RoboCasa365, compared with a tool-calling baseline using the same GPT-6 Astra planner and the same underlying robot primitives, it raises overall success from 63.1% to 71.7% under equal LLM-call budgets and, on instances solved by both agents, uses 49% fewer LLM calls and 65% fewer input tokens.

Source-provided article image: Fewer Tokens, Better Action: GPT-6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens
Figure 2 ·

Figure 2 : PyRUA-Lean combines primitive composition with selective feedback. (a) The tool-calling agent invokes robot primitives through tool calls. Operations that depend on preceding results generally require another VLM turn, and motion calls automatically return images and state. (b) PyRUA-Lean instead generates Python cells that compose robot primitives with helper functions, conditional checks, and local retries. Intermediate execution stays within the runtime, while explicitly requested images and state messages are recorded and returned at cell end for replanning. (c) A recorded placement example illustrates both mechanisms: one PyRUA-Lean cell replaces baseline steps 14–17, reducing four LLM turns to one. The cell computes a placement target from scene geometry (compose), conditionally executes lowering and release (compose), and requests an image only if the task remains unfinished (select). This successful execution returns five printed lines and no image, demonstrating how programmatic composition and selective feedback reduce repeated model interaction. Appendix A follows the whole episode.

arXiv

Interpretation

Under equal LLM-call budgets, PyRUA-Lean raises overall task success from 63.1% to 71.7%. Prior VLM agents control robots through visual feedback and action primitives but incur substantial token overhead from repeated model invocations and redundant observations; this work couples feedback-driven primitive composition with selective observation inside an interactive code-execution framework. Evidence comes from 700 task instances across three simulated benchmarks, LIBERO-PRO, RoboTwin 2.0, and RoboCasa365, compared against a tool-calling baseline using the same GPT-6 Astra planner and the same underlying robot primitives under equal LLM-call budgets.

On instances solved by both agents, PyRUA-Lean uses 49% fewer LLM calls and 65% fewer input tokens. The efficiency gain comes from having the agent compose classical robot primitives and learned VLA policies into Python cells that perform conditional checks and local retries, returning only explicitly requested images and state feedback for replanning. This efficiency comparison is restricted to instances both agents solved, making it a paired comparison on a shared task subset rather than over all 700 instances.

The framework establishes code execution with selective observation as a workable route to lower token overhead in VLM robot agents. Compared with turn-by-turn tool calling, moving control logic into executable code cells lets conditional checks and retries happen locally, reducing round trips to the model. Evidence is a controlled comparison on three simulated robot benchmarks, with the planner and underlying primitives held constant across both conditions, so the difference is attributable to the interaction framework itself.

Perspective

The result targets robot-agent settings that use a VLM planner and have composable classical primitives plus learned VLA policies, especially multi-step manipulation tasks that must be completed under a limited LLM-call budget. It shows that on simulated benchmarks, placing control logic in executable code cells and retrieving observations only on request can improve both success and call efficiency at equal budget, which is directly relevant to researchers and engineers building low-cost robot-agent pipelines.

This document is abstract-level material without the paper body, figures, or ablation details, so it is not possible to tell how the success gain distributes across the three benchmarks, which task types benefit most, or how much selective observation and local retries each contribute. In addition, all evidence comes from simulated task instances, so behavior on real robots and transferability across different planners or primitive libraries remain open questions to watch.

Sources