In-Context Robot Learning with General-Purpose VLM Agents: The GPT-Policy Framework
Synopsis
This work introduces GPT-Policy, a framework that connects an off-the-shelf general-purpose vision-language model (VLM) to robot tools through a shared context-to-action closed loop, and evaluates its in-context learning on real robots across five context families (human videos, robot videos with actions, goal images, self-interaction history, and online human-robot interaction), finding that task-relevant context can improve success while reducing decisions and execution time, yet better task understanding does not ensure precise contact, reliable outcome verification, or physical safety.
Figure 1 : In-context robot control with GPT-Policy. A general-purpose VLM combines the task instruction, initial state, and contextual information to guide robot actions. Context can include human videos, robot videos with recorded actions, goal images, human–robot interaction, and self-interaction history. The examples illustrate manipulation, interactive play, and mobile object retrieval.
arXivInterpretation
It proposes GPT-Policy, which links a fixed-parameter general-purpose VLM to robot tools through a shared closed-loop interface, letting the model select parameterized tool actions from context at test time without gradient updates. Unlike prior in-context robot learning work centered on dedicated robot policies or embodied foundation models, this study examines the context-use capability of an off-the-shelf general multimodal agent without robot-specific training or test-time parameter updates. The method provides a full problem formulation (a_t ~ π_θ(·|T,c_t,o_t,f_{t-1})), context construction, a structured tool interface, pose interpolation (linear plus SLERP), inverse-kinematics residual constraints, and Ruckig-based timing, with appendix configuration details for the YAM, ARX X5, and Morphi Kino platforms.
On real robots, a single human demonstration video improves task completion without robot action labels: Pick Red Towel and Pick Up Notebook both rise from 0/3 under None to 2/3 under Human Video, with lower average decision counts and execution times. This is consistent with transferring an interaction strategy across embodiments and indicates that action labels are not a prerequisite for in-context learning, complementing prior demonstration-conditioned control that relies on robot sensorimotor trajectories or geometric demonstration representations. Each task is repeated three times with task-specific geometric and semantic success criteria; towel pickup drops from 96.3 to 76.7 decisions and 24.6 to 18.9 minutes, notebook pickup from 94.0 to 66.7 decisions and 24.6 to 16.1 minutes, though the sample is only three trials.
On contact-sensitive tasks, time-aligned robot action references yield the highest observed success: bottle opening goes 0/3, 2/3, 3/3 and plug reinsertion 0/3, 0/3, 2/3 across None, Robot Video, and Robot Video + Action. Relative to sparse keyframe video alone, action references supply intermediate commanded poses and gripper transitions that reduce ambiguity about motion between keyframes, producing closer alignment with demonstrated orientations and support posture. Bottle opening retains 205 action samples for 13 video keyframes and plug reinsertion 131 samples for 14 keyframes; success improves while decision counts and times do not consistently fall, and averages include failed trials.
Under goal images, self-interaction history, and online human-robot interaction, GPT-6 Astra reaches 3/3 success on the corresponding tasks and exhibits goal grounding, changed operation order, and online coordination. These cases show that a single goal image can jointly specify object identity, relative position, and spacing, that self-interaction history supports intermediate-subgoal reasoning such as removing a towel to uncover the pink plate, and that online interaction history supports tracking turn-taking and pointing intent. Each task is repeated three times; self-history tasks average 35.3 decisions and 8.1 minutes (Lemon to Pink Plate) and 40.33 decisions and 25.53 minutes (Movable Exploration); human-robot interaction averages 69.7 decisions and 13.6 minutes (Tic-Tac-Toe) and 67.3 decisions and 15.0 minutes (Pointed Fruit Pickup).
Perspective
The work targets small task series under selected platform, model, and context conditions, and is meant for studying how a general VLM uses demonstrations, goal images, and interaction experience during closed-loop robotic-arm operation; its shared context-to-action interface and embodiment adapters can be reused on further platforms and tasks and serve as a basis for testing when context improves success, efficiency, and recovery.
A careful reader would still watch whether the success and efficiency gains from context hold over larger task series and more platforms; why action references raise success without consistently lowering cost; how safety events such as inter-arm collisions are systematically monitored and reported given that collision checking is not part of the Cartesian planner; and how the gap between model-declared completion and physical task success is independently verified. In addition, this is a fast parse of the text in which figures (Figures 4 to 8) and some table details are not fully rendered, which may limit further judgment about the qualitative comparisons and alignment metrics.
