Public articles linked to the same research event.
arXiv PyRUA-Lean is an interactive code-execution framework in which a VLM agent composes classical robot primitives and frozen VLA policies into Python cells that perform conditional checks and local retries and return only explicitly requested images and state feedback; across 700 simulated task instances from LIBERO-PRO, RoboTwin 2.0, and RoboCasa365, and against a tool-calling baseline using the same GPT-6 Astra planner and the same primitives, it raises overall success from 63.1% to 71.7% under equal LLM-call budgets and uses 49% fewer LLM calls and 65% fewer input tokens on instances solved by both agents.
PyRUA-Lean is an interactive code-execution framework in which a VLM agent composes classical robot primitives and frozen VLA policies into Python cells that perform conditional checks and local retries and return only explicitly requested images and state feedback; across 700 simulated task instances from LIBERO-PRO, RoboTwin 2.0, and RoboCasa365, and against a tool-calling baseline using the same GPT-6 Astra planner and the same primitives, it raises overall success from 63.1% to 71.7% under equal LLM-call budgets and uses 49% fewer LLM calls and 65% fewer input tokens on instances solved by both agents.
PyRUA-Lean is an interactive code-execution framework in which a VLM agent composes classical robot primitives and frozen VLA policies into Python cells that perform conditional checks and local retries and return only explicitly requested images and state feedback; across 700 simulated task instances from LIBERO-PRO, RoboTwin 2.0, and RoboCasa365, and against a tool-calling baseline using the same GPT-6 Astra planner and the same primitives, it raises overall success from 63.1% to 71.7% under equal LLM-call budgets and uses 49% fewer LLM calls and 65% fewer input tokens on instances solved by both agents.
PyRUA-Lean is an interactive code-execution framework in which a VLM agent composes classical robot primitives and frozen VLA policies into Python cells that perform conditional checks and local retries and return only explicitly requested images and state feedback; across 700 simulated task instances from LIBERO-PRO, RoboTwin 2.0, and RoboCasa365, and against a tool-calling baseline using the same GPT-6 Astra planner and the same primitives, it raises overall success from 63.1% to 71.7% under equal LLM-call budgets and uses 49% fewer LLM calls and 65% fewer input tokens on instances solved by both agents.
arXiv The authors introduce PyRUA-Lean, an interactive code-execution framework that couples feedback-driven primitive composition with selective observation, letting the agent compose classical robot primitives and learned vision-language-action (VLA) policies into Python cells that perform conditional checks and local retries and return only explicitly requested images and state feedback for replanning; across 700 simulated task instances from LIBERO-PRO, RoboTwin 2.0, and RoboCasa365, compared with a tool-calling baseline using the same GPT-6 Astra planner and the same underlying robot primitives, it raises overall success from 63.1% to 71.7% under equal LLM-call budgets and, on instances solved by both agents, uses 49% fewer LLM calls and 65% fewer input tokens.
The authors introduce PyRUA-Lean, an interactive code-execution framework that couples feedback-driven primitive composition with selective observation, letting the agent compose classical robot primitives and learned vision-language-action (VLA) policies into Python cells that perform conditional checks and local retries and return only explicitly requested images and state feedback for replanning; across 700 simulated task instances from LIBERO-PRO, RoboTwin 2.0, and RoboCasa365, compared with a tool-calling baseline using the same GPT-6 Astra planner and the same underlying robot primitives, it raises overall success from 63.1% to 71.7% under equal LLM-call budgets and, on instances solved by both agents, uses 49% fewer LLM calls and 65% fewer input tokens.
The authors introduce PyRUA-Lean, an interactive code-execution framework that couples feedback-driven primitive composition with selective observation, letting the agent compose classical robot primitives and learned vision-language-action (VLA) policies into Python cells that perform conditional checks and local retries and return only explicitly requested images and state feedback for replanning; across 700 simulated task instances from LIBERO-PRO, RoboTwin 2.0, and RoboCasa365, compared with a tool-calling baseline using the same GPT-6 Astra planner and the same underlying robot primitives, it raises overall success from 63.1% to 71.7% under equal LLM-call budgets and, on instances solved by both agents, uses 49% fewer LLM calls and 65% fewer input tokens.
The authors introduce PyRUA-Lean, an interactive code-execution framework that couples feedback-driven primitive composition with selective observation, letting the agent compose classical robot primitives and learned vision-language-action (VLA) policies into Python cells that perform conditional checks and local retries and return only explicitly requested images and state feedback for replanning; across 700 simulated task instances from LIBERO-PRO, RoboTwin 2.0, and RoboCasa365, compared with a tool-calling baseline using the same GPT-6 Astra planner and the same underlying robot primitives, it raises overall success from 63.1% to 71.7% under equal LLM-call budgets and, on instances solved by both agents, uses 49% fewer LLM calls and 65% fewer input tokens.