Rewriting the Instruction Lifts Robot Success: Rules Distilled from Eight Tasks Raise a Frozen VLA by 16–27% Relative on Twelve Held-Out Tasks
Synopsis
This work characterizes the language sensitivity of vision-language-action models (VLAs): one-word edits can move success by tens of points, and phrasing alone nearly closes the 21-point gap between in-distribution and out-of-distribution tasks; the authors have a large language model distill scored phrasings of a few training tasks into ten to twenty rephrasing rules, rewrite each incoming instruction once at deployment, and thereby improve the frozen policy by 16–27% relative on twelve held-out tasks without modifying its weights, lifting in-finetune success on LIBERO from 93.6% to 97.8%.
Fig. 1 : Method overview. Evidence gathering : a VLM rephrases each training task, and the frozen VLA’s rollout success on each phrasing forms the evidence 𝒟 \mathcal{D} . Rule distillation : an LLM distills 𝒟 \mathcal{D} , once and offline, into 10–20 explicit phrasing rules R R . Rule application : at deployment a VLM conditioned on R R rewrites each incoming instruction once ( f R f_{R} ) before passing it to the unchanged VLA. The phrasings, scores, and rules shown are illustrative only.
arXivInterpretation
VLA language sensitivity is systematic and can be characterized by statistically tested single-edit swings: “switch on the stove” succeeds 100% of the time while “switch on the hot plate” succeeds 2%, and on a SIMPLER task “purple eggplant goes on the sponge” succeeds 69% while dropping the color adjective lowers it to 8%. Prior work such as LIBERO-PRO and LIBERO-Plus reported instruction perturbations costing 14 to 50 points but largely attributed this to models ignoring the instruction; this work offers the complementary, finer-grained observation that when the policy does attend to language, a single word can move success by tens of points, and it retains only statistically significant swing pairs. On SIMPLER/Bridge each phrase is evaluated 72 times on 24 fixed scene layouts with 3 repetitions each, and on LIBERO once on each of 50 layouts; significance uses a two-proportion z test, and sets whose gap could reflect scene ambiguity are manually removed (eight of 23 significant sets on LIBERO).
Phrasing alone leaves substantial headroom: an oracle phrase search shows that wording nearly closes the 21-point gap between in-distribution and out-of-distribution tasks. This quantifies that much of what appears to be a generalization gap is attributable to phrasing rather than genuine task novelty; the authors run three rounds of VLM generation and scoring to select the best phrase and re-measure it on held-out layouts to confirm selection does not inflate the headroom. On SIMPLER candidates are screened on layouts 0–17 and the selected phrase is re-measured 12 times per layout on all 24 layouts, with held-out layouts 18–23 alone giving the same headroom; on LIBERO candidates are screened on 5–10 layouts and the oracle bar pools each winner’s re-measurement on disjoint layouts.
Expressing the sensitivity as explicit rephrasing rules reduces it without modifying the policy: ten to twenty rules distilled from eight evidence tasks, applied once per instruction at deployment, improve the frozen policy by 16–27% relative on twelve held-out tasks, with gains concentrated on out-of-distribution tasks. Unlike test-time verification such as CoVer, which selects among 40 phrase-action candidates at every step, and unlike per-task instruction optimization, this method acts on phrasing only, never on actions, rephrases once per episode so per-step inference cost is unchanged, and transfers zero-shot to tasks that contributed no evidence. Evaluated on 12 sealed Bridge tasks and 24 fixed layouts across 72 adversarial, 363 human-generated, and 186 VLM-generated phrases; effects use a paired sign-flip permutation test and are stable across three applier models (Claude, Gemini, Qwen) and three independent distillations, with pooled success varying by 0.7 to 5.1 points across draws.
The pipeline replicates on LIBERO: three independently distilled rulebooks lift in-finetune success from 93.6% to 97.8% on average while leaving out-of-finetune success unchanged, and the no-rules rephraser gains nothing in-finetune. The replication uses a preregistered sealed set of 20 tasks (10 in-finetune, 10 out-of-finetune) with 10 VLM-generated natural rephrasings per task and evidence from non-sealed single-edit pairs only, indicating the rulebook effect is not tied to one benchmark or one applier model. Each arm is evaluated over 50 layouts per arm; gains are attributable to specific repairs, such as rewriting hot plate, hotplate, or burner to stove gaining 45 points on roughly 9 of the 200 test phrases, rewriting dish to plate gaining 24 points, and rewriting onto to on and into to in gaining 2.5 points over roughly 50 phrases.
Perspective
The results apply to frozen VLAs evaluated in simulation: WidowX tabletop tasks in SIMPLER/Bridge and LIBERO tabletop tasks, with success determined procedurally by the benchmark. The method suits deployment settings where scored phrasing evidence exists and where one wants to improve how a frozen policy parses natural-language instructions without retraining it; once distilled, a rulebook can be reused offline and applied zero-shot to tasks and instructions that contributed no evidence. The authors note that rulebooks distilled from more tasks may yield larger gains, and that real robots are a natural next step.
The canonical set has only one phrase per task, and the authors explicitly state it is statistically underpowered and should be read with caution. The in-finetune rulebooks underperformed, which the authors suggest may relate to a weak proxy score, whereas the LIBERO rulebooks, whose evidence was entirely rollout-scored, did not. The single-edit search applies no multiple-comparison correction, and the authors note that with 426 candidate pairs about 21 significant swings are expected by chance alone, with manual removal of sets that could reflect scene ambiguity. Evaluation is in simulation only, so real-robot behavior remains open, and the authors list systematic cross-model measurement of language sensitivity and attributing it to a cause as the most valuable next direction.
