RL post-training of vision-language-action models concentrates parameter updates in Timestep Modules that hold only 14.83%–27.58% of the action expert, and their low-rank directions predict and improve task success
Synopsis
This work systematically analyzes how reinforcement learning (RL) reshapes flow-based vision-language-action (VLA) models, finding that across π0.5 and GR00T N1.5/N1.6 on LIBERO, ManiSkill, MetaWorld, and CALVIN, RL induces low-rank parameter updates highly concentrated in the action expert's Timestep Modules, which hold 14.83%–27.58% of the parameters yet capture a disproportionate share of the RL gain, with shift-vector update directions predicting task success at up to 99.6% ROC-AUC and steering along them improving policies without additional RL training.
Interpretation
RL updates to VLAs are highly concentrated in the Timestep Modules and exhibit a strongly low-rank structure. Parameter-space understanding of VLA RL was largely unexplored, whereas LLM RL has studies on sparse updates and low-rank directions; this work extends that lens to flow-based VLA action experts and identifies the Timestep Modules, typically omitted from standard LoRA configurations, as the dominant update site. Update density and effective rank are analyzed across π0.5, GR00T N1.5/N1.6, SmolVLA, and LIBERO, ManiSkill, MetaWorld, CALVIN; RL Timestep Module updates have effective rank of only about 1–4, roughly an order of magnitude lower than BC, while RL-trained image/video DiTs do not show the same low-rank structure.
Timestep Modules capture a disproportionate share of the RL performance gain. Through module-replacement experiments, replacing attention and FFN parameters with their base versions while keeping only RL-trained Timestep Modules preserves much of the RL gain, whereas the reverse replacement drops performance; LoRA targeted at Timestep Modules converges faster and reaches higher success than standard LoRA. Controlled module replacement on π0.5, GR00T N1.5/N1.6, where Timestep Modules comprise 14.83%–27.58% of action expert parameters; reconstructing updates from only the top four singular directions retains most performance, while removing them largely degrades it.
RL specializes Timestep Modules to the discrete denoising timesteps used during rollouts, and among their outputs the shift vector changes most distinctly and encodes task-outcome information. Standard BC trains on continuous timesteps and shows smooth responses, whereas RL and discrete-timestep BC both produce localized responses and low-rank updates, indicating discrete-timestep training is the main cause; among scale, shift, and gate, the shift vector's RL-induced change is least aligned with its base direction, introducing the most new directions. SVD of AdaRMS updates and cosine similarity to per-timestep conditioning vectors; probing with shift update directions reaches ROC-AUCs of 96.6%–99.6% on LIBERO-Spatial, LIBERO-Object, LIBERO-Goal, and ManiSkill, with single directions also performing well and random-label controls clearly lower.
The geometry of shift updates reflects task relationships and can be used for inference-time steering to improve an already trained policy. Pairwise similarity of shift update directions correlates positively with cross-task transfer patterns, offering a compact representation of task relationships; adaptive steering along shift update directions improves the final RL policy without additional RL training and outperforms matched-norm random directions. Single-task RL policies are trained for ten LIBERO-Spatial tasks; mean task-wise Spearman correlation between shift-update alignment and transfer effects is positive (range includes negative values); steering improves success by roughly 1.3–3.3 points across five benchmarks, e.g., LIBERO-Spatial from 86.7% to 90.0% and LIBERO-Goal from 90.7% to 94.7%.
Perspective
The results target VLA models with flow-based action experts, validated on simulation benchmarks (LIBERO, ManiSkill, MetaWorld, CALVIN) and checkpoints including π0.5, GR00T N1.5/N1.6, and SmolVLA, mostly with PPO (GRPO for SmolVLA). They are meant for researchers and engineers who want to perform RL post-training with fewer parameters or use Timestep Modules and shift directions for lightweight adaptation and inference-time steering; discrete-timestep BC and residual RL design are noted as follow-up directions.
The low-rank structure does not consistently appear in RL-trained image/video DiTs, so whether it is specific to action policies needs more architectural validation; probing and steering AUCs and gains fluctuate across benchmarks such as LIBERO-Goal and MetaWorld, and the sign of shift directions varies across sublayer–timestep pairs; the cross-task transfer correlation range includes negative values, so the robustness of the task-relationship representation remains to be extended; additionally, the trade-off between discrete- and continuous-timestep BC and the choice of steering coefficient still need testing in more real-robot settings.
