StructRL lifts long-horizon VLA success from 41.5% to 49.1% with verifiable subtask rewards
Synopsis
StructRL is an online reinforcement learning framework that automatically decomposes each long-horizon instruction into subtasks verifiable by binary environment-state criteria and organizes them into ordered dependency groups, granting intermediate rewards only after all prerequisites of a subtask are complete and scaling each reward by completion pace; across RoboCasa365 and LIBERO-Long with the GR00T-N1.5 and π0.5 VLA backbones it consistently outperforms the evaluated online RL baselines (e.g., on GR00T-N1.5, RoboCasa365 rises from 41.5% to 49.1% and LIBERO-Long from 92.4% to 96.6%).
Interpretation
StructRL decomposes long-horizon commands into verifiable subtasks with a dependency structure, and uses structure-aware reward gating so that a detected completion earns intermediate credit only after all of its prerequisites are satisfied, with each subtask contributing reward at most once. Prior online VLA RL largely relies on a terminal reward issued only when the whole task succeeds, or on learned scalar progress estimates; the former cannot distinguish early failures from near-success rollouts, and the latter does not specify which prerequisite events make a detected completion valid progress. StructRL encodes both whether an event occurred and whether it constitutes valid progress given the events completed so far. The paper gives a formal definition (prerequisite sets and a gating indicator) and shows decompositions for the 16 RoboCasa365 kitchen tasks, e.g., in Pack Lunch, closing the box counts as valid progress only after both items have been placed.
StructRL uses dynamic reward pacing to scale reward magnitude by completion pace: it references the average interval in SFT demonstrations from when a subtask's prerequisites first become complete to when that subtask is completed, giving larger rewards for faster completions, with a scale parameter bounding each subtask reward. Fixed-magnitude subtask rewards cannot distinguish direct completions from delayed ones involving unnecessary wandering, which blurs credit assignment in long-horizon rollouts; StructRL brings demonstration-derived durations into the reward magnitude. The ablation shows that on GR00T-N1.5, adding subtask rewards first raises success from 41.3% to 47.4% on RoboCasa365 and from 91.2% to 94.7% on LIBERO-Long; dynamic pacing adds 0.5 and 1.2 points, and structure-aware gating adds another 1.3 and 0.5 points, reaching 49.2% and 96.4%.
Across two benchmarks and two VLA backbones, StructRL consistently exceeds the evaluated online RL baselines, with gains concentrated in the longer-horizon task buckets. The paper compares under the same SFT initialization and the same environment-interaction budget against Sparse-RL, SimpleVLA-RL, and PolicyTrim, and additionally against the learned dense-reward baseline Robometer under a matched PPO setup. With GR00T-N1.5, overall success is 49.1% versus 41.5% for the strongest online baseline on RoboCasa365 (+7.6 points) and 96.6% versus 92.4% on LIBERO-Long (+4.2 points); with π0.5 the gains are 3.9 and 2.2 points; the largest bucket gains are +7.0 points on the RoboCasa365 1400–2900-step bucket and +6.6 points on the LIBERO-Long 340–400-step bucket. It leads Robometer by 2.4 and 3.2 points respectively.
The structured reward is compatible with GRPO and remains mildly positive on shorter-horizon tasks, while zero-shot transfer to unseen task compositions remains limited. The paper tests the reward formulation's portability across optimizers and its applicability where baselines are already near saturation, and reports transfer and forgetting on tasks excluded from RL training. The GRPO variant beats SimpleVLA-RL on RoboCasa365 by 3.4 points (GR00T-N1.5) and 1.3 points (π0.5) but trails the PPO variant by 4.2 and 2.6 points; on LIBERO Spatial/Object/Goal the average rises from 94.7% to 95.7%; zero-shot Composite-Unseen reaches only 4.8% (comparison 4.3%), and Atomic-Seen reaches 20.5% (SFT 17.0%, SimpleVLA-RL 16.9%).
Perspective
The result targets post-training of long-horizon manipulation policies after supervised fine-tuning on demonstrations, in settings that can supply binary-verifiable subtask completion signals, such as the RoboCasa365 kitchen composite tasks and LIBERO-Long tabletop tasks; for VLA post-training pipelines that want dense supervision with less reward engineering, it offers a path that does not require reward-model inference during rollout collection. The paper also reports that the reward form works with GRPO and stays mildly positive on shorter LIBERO Spatial, Object, and Goal suites, so its scope is not limited to the longest-horizon settings.
Decompositions are LLM-generated, and the paper reports that swapping to the weaker Qwen3.5-9B lowers overall success from 49.1% to 45.7%, indicating decomposition quality affects downstream RL performance; the density sweep peaks at 50.4% around an average of 5.0 subtasks and falls to 37.6% at 52.0, below the 40.2% of terminal-reward-only, which the paper attributes both to earlier, less task-aligned completion signals and to growth in total intermediate reward. Reward pacing relies on demonstration durations, and replacing them with uniform horizon allocation lowers success from 49.2% to 43.1%, leaving more adaptive demonstration-free estimates open. The ethics statement notes that some rollouts collected intermediate rewards and then idled until timeout, and that favoring faster subtask completion would require separate safety validation before deployment on physical robots. Zero-shot transfer to unseen task compositions remains low (4.8%), and real-world evaluation plus deriving completion signals from perception are listed as future directions.
