Skip to main content
Back to timeline
arXivSource publication:

RoboFollow diagnoses nine embodied policies in high-entropy scenes: near-saturated L0, sharp L1–L3 intent drops, and no fix from stronger VLMs or existing optimizations

Synopsis

The work introduces RoboFollow, a diagnostic benchmark combining high scene entropy, an L0–L3 hierarchical perturbation protocol, and stage-wise Intent/Execution scoring, and finds across nine VLA and WAM policies that strong in-distribution L0 performance does not transfer to L1–L3, with stronger VLM backbones, QA co-training, LangForce, and classifier-free guidance all failing to close the gap.

AI-generated editorial illustration: RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents

Interpretation

RoboFollow turns "is language necessary" into a measurable dataset property: the same initial scene configuration supports multiple kinematically feasible and semantically valid task branches, so vision alone cannot identify the intended behavior. Existing manipulation benchmarks emphasize final task success, and many episodes contain a dominant behavior inferable from the initial observation, making language partly redundant; RoboFollow quantifies this via scene entropy, the conditional entropy of training task labels given the task-independent initial scene specification, and reports values above the equally weighted LIBERO Spatial/Object/Goal/Long suites under the same task-label definition. The paper gives a formal definition and grouped computation of scene entropy, with 3,750 training episodes across four scenes and 50 demonstrations per task; Scenes 1/2/3 each contain one group of 16 task labels, and Scene 4 contains three groups of 16, 7, and 4.

The L0–L3 protocol with stage-wise Intent and Execution scores separates semantic misunderstanding from motor failure, and the results point to semantic intent selection rather than low-level control as the main source of failure. Conventional binary final-state success penalizes policies that pick the right intent but execute imperfectly, and rewards policies that reach the final state through semantically incorrect intermediate actions; RoboFollow reports Intent and Execution Scores per stage and includes a finishing stage worth 20% that requires refraining from further unrelated actions after task completion. On Scenes 1 and 2, some models reach near-saturated L0 intent scores, reported as 99.1% and 100.0%, with another at 98.2% and 95.6%; yet in Scene 1 one model falls from 98.2% at L0 to 0.0% at L1, another from 99.1% to 45.5% at L1 and 44.9% at L3, and in Scene 2 one model drops from 100.0% at L0 to 34.2% at L3.

A mechanistic diagnosis on Scene 2 reveals two cascading bottlenecks: the VLM backbone itself often lacks scene understanding, and even when comprehension succeeds, the action generation pipeline does not faithfully translate it into behavior. The paper uses a 20-question visual QA probe to independently test VLM scene comprehension and the fidelity with which comprehended semantics propagate to action generation, rather than inferring causes only from end-to-end success rates. Base PaliGemma, the pre-trained backbone, and the fine-tuned backbone answer 1/20, 2/20, and 3/20 questions correctly; for questions the fine-tuned VLM answers correctly, about half of those episodes still result in incorrect manipulation behavior; a comparable disconnect between correct visual comprehension and successful manipulation execution is observed in Motus, whose VLM backbone is entirely frozen.

Three representative mitigation strategies all fail to push generalization beyond L0 in the tested configuration: a stronger VLM backbone with QA co-training, LangForce, and classifier-free guidance. The paper places three common routes—upgrading the backbone, preserving linguistic competence during training, and amplifying language conditioning at inference—under the same high-entropy diagnosis and characterizes how each falls short. Qwen3-VL-4B reaches 19/20 on the fine-grained Scene 2 QA probe, but combining it with the GR00T diffusion action head and QA co-training still yields a collapse in out-of-distribution instruction following across L1–L3; LangForce yields only marginal changes; CFG at guidance scales 1.2 and 1.5 degrades even L0 performance and produces erratic trajectories.

Perspective

The benchmark targets embodied-policy development and evaluation settings where semantic misselection must be distinguished from poor motor execution, and it applies to short-horizon manipulation in controlled high-entropy simulated scenes whose design principle is to make language necessary rather than optional. For users, it supports introducing task ambiguity during training, separating visual grounding, semantic recombination, and joint generalization via L0–L3 at evaluation time, and localizing failure stages through Intent and Execution Scores. The paper also notes that its controls vary demonstration count, instruction variants, and training duration, but do not test increased structural layout diversity during training, since that would require new training layouts while keeping evaluation layouts held out.

The paper explicitly positions RoboFollow as a diagnostic benchmark whose controlled scenes simplify object geometry, task horizons, and interaction dynamics, so it does not cover long-horizon planning, contact-rich manipulation, open-vocabulary diversity, or large-scale real-world deployment. The real-robot pilot compares different instruction sets and pick/stack compositions rather than matched task pairs, and the paper states that transfer across simulators or to more complex real-world tasks has not been established. Fine-tuning controls show that more demonstrations, instruction variants, and training steps improve some Scene 2 metrics, but L0–L2/L3 gaps remain substantial, and structural layout diversity is untested. In addition, the loaded text is the full paper, but per-model numbers in the tables appear as prose descriptions, so exact per-model, per-level scores would still require consulting the original tables.

Sources