ART turns VLA models into tool-calling robot agents, reporting a 20% higher success rate than mainstream baselines in simulation and real-world tasks
Synopsis
The work proposes Agentic Robot with Tool-use (ART), a tool-injection framework that tunes any VLA model to leverage off-the-shelf tool modules for low-level vision, high-level affordance, and embodiment enhancement; the authors built a dataset of 30K tool-use trajectories and action demonstrations and designed a training regimen for long-trajectory tool-use reasoning in challenging environments, with experiments showing ART achieves a 20% higher success rate than mainstream baselines on simulation and real-world tasks such as pick-and-place in the dark at novel viewpoints.
Figure 1 : Comparison of VLA Paradigms , including (a) standard end-to-end VLA [ 28 , 4 , 5 ] , (b) VLA enhanced with Chain-of-Thought reasoning [ 50 ] , (c) modular robotic behavior synthesis [ 31 , 34 ] , and (d) our Agentic Robot with Tool use (ART) framework, which intDegrates tool use into the end-to-end action generation process. The ART model demonstrates enhanced adaptability to complex environments (Task 1) and correct hallucinated predictions in foundation models (Task 2) through dynamic tool use.
arXivInterpretation
ART integrates end-to-end VLA models with agentic tool-use, reducing the complexity of the action solution space through tool-use, which improves generalizability across tasks and reduces data dependency. Compared with vanilla VLA models that use a whole continuous action solution space, ART rewrites the action-solving path via tool injection, delegating low-level vision, high-level affordance, and embodiment enhancement to off-the-shelf tool modules. The abstract argues at the framework level and names generalizability and low data dependency as design goals; mechanism details are not expanded in the provided text.
The authors built a dataset of 30K tool-use trajectories and action demonstrations, much smaller than those used by baseline methods. The dataset was constructed specifically to demonstrate ART's low data dependency rather than reusing existing large-scale data. The abstract states the 30K figure and that it is 'much smaller than those used by baseline methods'; no per-baseline data-size comparison is given.
The authors designed a training regimen for long-trajectory tool-use reasoning in challenging environments. The regimen targets long-trajectory tool-use reasoning rather than single-step action prediction, supporting multi-step tool calls in complex environments. The abstract only states the regimen's existence and target setting, without training details, stages, or hyperparameters.
Experiments show ART achieves a 20% higher success rate than mainstream baselines on simulation and real-world tasks, including pick-and-place in the dark at novel viewpoints. Relative to mainstream baselines, ART reports success-rate gains in both simulation and real-world settings, covering perceptually adverse conditions such as darkness and novel viewpoints. The abstract reports the 20% success-rate improvement and names the task types; no sample sizes, confidence intervals, or per-task breakdown are provided.
Perspective
The framework targets robot manipulation tasks requiring low-level vision, high-level affordance, and embodiment enhancement, applies to VLA models that can be connected to off-the-shelf tool modules, and is validated on pick-and-place-style tasks in simulation and the real world, including dark and novel-viewpoint conditions. For teams wanting to train with less data and later add new tools incrementally, ART offers a modular path.
The provided text is the arXiv abstract page, lacking the body, figures, and experimental tables, so the specific tool-module interfaces, the staged design of the long-trajectory training regimen, the statistical basis of the 20% gain, and per-task results cannot be confirmed here. Readers should still watch for fallback behavior when tool calls fail, the extra tuning cost of adding new tools, and how the framework performs on tasks and embodiments not listed in the abstract.
