RoboDawn lets a frozen VLM control robots zero-shot through discrete commands, with one demonstration lifting RoboTwin 2.0 C2R success from 53.2% to 73.6%
Synopsis
The work introduces RoboDawn, which exposes robotic control to a frozen VLM as a compact set of discrete translation, rotation, and gripper commands in a closed observe-reason-act-reobserve loop, paired with an in-context learning scheme; it reaches 53.2% zero-shot and 73.6% one-shot success on RoboTwin 2.0 C2R, improves RoboDojo from 35.67% to 47.17%, and transfers zero-shot to a Franka robot for block-in-basket and block stacking.
Interpretation
RoboDawn formulates manipulation as closed-loop visual decision making: the VLM repeatedly observes the current visual state, reasons about the next action, executes it, and adapts subsequent decisions to the resulting state, using discrete semantic primitives for translation, rotation, and gripper rather than low-level continuous control. Unlike VLA and world-action models that learn observation-to-action mappings from large robot datasets, this approach keeps VLM parameters frozen and instead relies on a human-intuitive action interface resembling abstractions from games and teleoperation. The paper specifies the full command grammar, notes that translation and rotation magnitudes are clipped, defines poses relative to the gripper interaction point, and states that online control does not rely on privileged object poses; evaluation covers 50 bimanual tasks in RoboTwin 2.0 C2R with 10 independent runs per task.
An interface-aligned in-context learning scheme uses a few demonstrations to ground both primitive semantics and task-level strategies without parameter updates, and a single demonstration already yields substantial gains. Demonstrations are re-expressed in the semantic command space and split into a shared command primer and task-level demonstrations, with sparsification keeping the in-context image budget to 16 observations per round for long-horizon trajectories. On RoboTwin 2.0 C2R, GPT-6 Astra rises from 53.2% zero-shot to 73.6% one-shot; on RoboDojo from 35.67% to 47.17%; ablations show Gemini-3.8-Flash going from 47.0% to 62.2%, 65.4% with four demonstrations, and 62.7% with eight.
Performance scales with the underlying VLM and with test-time command budget, showing test-time scaling behavior. This suggests that under a fixed interface, stronger models translate directly into higher success, and that additional interaction compute can be converted into task success rather than relying only on more robot data. Under one-shot, success rises across models from 14.4% with GPT-5.6-Luna, 43.2% with GPT-5.6-Sol, 62.2% with Gemini-3.8-Flash, to 73.6% with GPT-6 Astra; on RoboDojo the one-shot rate rises from 31.2% at 60 commands to 47.2% at 240 commands, and zero-shot from 23.7% to 35.7%.
The same framework transfers zero-shot to real robots, though task difficulty varies markedly. The real-robot results indicate the interface is not confined to simulation while highlighting that rotation-heavy and precision-placement tasks remain harder. On Franka, block-in-basket reaches 9/10 and block stacking 5/10, while cloth folding on Piper reaches 0/10; the authors attribute the cloth-folding difficulty to substantial end-effector rotation and orientation adjustment and note several near-successful trials whose final folds were not neat enough to count.
Perspective
The work targets high-level manipulation decisions: the authors explicitly leave low-level dexterity and high-frequency control to other directions and position RoboDawn as a high-level decision-making engine. It is meant for researchers and engineering teams who want to use an existing VLM for manipulation without task-specific parameter updates, in settings such as RoboTwin 2.0 C2R's clean-to-randomized generalization, RoboDojo's generalist manipulation evaluation, and Franka block-in-basket and block stacking. The in-context scheme assumes no fixed number of demonstrations, so zero-shot, one-shot, and few-shot all map onto the same interface and the demonstration budget can be tuned to task difficulty.
Inference efficiency is an explicit open question: RoboDawn shows 9.74 s inference, 3.4 commands on average, 2.09 s motion, and an inference-to-motion ratio of 4.65, which the authors say makes it less suitable for high-frequency control. Uneven difficulty across degrees of freedom also remains: the VLM handles translation more reliably than precise arm and end-effector rotations, and in-context learning alleviates but does not close the gap. Fine-grained interaction control is likewise limited, since the discrete semantic interface can be too coarse for subtle adjustments near contact, grasping, and placement. The 0/10 cloth-folding result and the authors' hypothesis that rotational motions are less represented in web-scale pretraining data mark this as an open direction. The RoboDojo failure analysis further lists three modes: insufficient manipulation precision, IK-related execution errors causing collisions, and a mismatch between the model's notion of task completion and the benchmark's success criterion. This summary is based on the full text and does not include every per-item numeric detail from the figures and tables.
