DRIVE turns successful-trajectory diversity into an RL objective, improving VLA out-of-domain generalization in simulation and on real robots
Synopsis
The authors analyze how reinforcement learning fine-tuning (RFT) reshapes exploration in vision-language-action (VLA) policies, finding that RFT contracts the overall behavior distribution while making successful trajectories more diverse, easier to elicit, and broader in task-valid solution coverage; they then propose DRIVE, which converts temporally aligned similarity among grouped rollouts under matched task conditions into a success-conditioned intrinsic reward, improving out-of-domain success on LIBERO-Plus, ManiSkill3, and RoboTwin 2.0 and raising average out-of-domain success from 64.1% to 73.3% on a dual-arm AgileX PiPER-X platform.
Figure 1: Beyond task success, diverse successful trajectories reveal broader solution modes. This diversity provides a new opportunity for improving VLA generalization.
arXivInterpretation
RFT produces two opposing trends at once: normalized GAK similarity rises across all rollouts but falls among successful ones, so global behavioral contraction masks successful-mode diversification. Prior work mainly examined performance gains from task-reward optimization; this analysis isolates successful-mode coverage as a lens for studying RFT. On ManiSkill3, SFT and intermediate/final RFT checkpoints are compared under matched task conditions with SDE sampling, measuring similarity separately over all and successful rollouts; success rates rise from 41% to 75% and from 52% to 77%.
RFT makes latent successful behaviors more accessible and covers a broader task-valid solution space under the same sampling budget. Pass@k measures how readily success can be elicited, while Coverage@k based on normalized GAK matching measures the breadth of successes, separating easier success from more ways to succeed. SFT and final RFT are evaluated on 32 task-initial-state configurations with up to 256 rollouts per policy and configuration; Coverage@k uses SFT successes as reference with leave-one-out matching to avoid trivial self-matches.
DRIVE converts temporally aligned similarity among grouped trajectories into a success-conditioned relative diversity potential, and potential-based reward shaping preserves optimal policy invariance. Unlike unsupervised skill discovery, policy diversification, or global behavioral variation, DRIVE compares trajectories only under matched task conditions and grants the diversity bonus only to successful trajectories, avoiding rewards for failures or purely timing differences. Trajectories use mean-pooled VLM-prefix representations already computed by the policy plus the terminal observation; GAK dynamic programming aggregates all admissible monotonic alignments with self-similarity normalization; Theorem 1 proves optimal policy invariance.
Across three simulation benchmarks and a real dual-arm platform, DRIVE improves out-of-domain success over Vanilla RFT. Under the same SFT initialization and a common PPO + Flow-SDE pipeline, DRIVE attains the highest benchmark-family macro-average on both backbones compared with sampling-, policy-, and optimization-level interventions such as Higher Noise, KL regularization, and Clip-Higher. Each backbone improves 11 of 12 benchmark-split results, with OOD macro-averages rising from 48.1% to 53.4% and from 71.1% to 73.1%; on real robots, average OOD success across two tasks and three distribution shifts rises from 64.1% to 73.3%, with 20 trials per task-condition pair.
Perspective
The results target robot manipulation settings where VLA policies are fine-tuned with RL under a PPO and Flow-SDE pipeline, and where rollouts can be grouped under matched task conditions. The authors note that extending DRIVE to longer-horizon tasks, broader embodiments, and online real-world fine-tuning would further assess its scalability and generality.
Pairwise GAK comparisons within groups add computation, and the diversity metric relies on VLM representations to capture meaningful behavioral differences. OOD performance fluctuates across training checkpoints rather than rising monotonically, so in-domain performance alone does not characterize generalization. Real-world evaluation covers two tasks and three distribution shifts with 20 trials per task-condition pair, leaving behavior on broader tasks and embodiments an open question.
