Skip to main content
Back to timeline
arXivSource publication:

SWAP routes stepwise among VLA policies via offline RL, lifting real-world DROID task success from 63.3% to 96.7%

Synopsis

SWAP casts robot policy routing as an offline reinforcement learning problem, training an IQL routing critic that picks, after each action chunk, the candidate pretrained VLA policy with the highest predicted long-horizon return given the current observation; on real-world DROID tabletop manipulation it raises success from 63.3% for the best single policy to 96.7% while cutting average action steps for successful trajectories from 269 to 193 (a 28.3% reduction), and it outperforms fixed-policy and random-routing baselines in LIBERO and LIBERO-Plus simulation.

Source-provided article image: SWAP: Stepwise Action Policy Routing for Vision-Language-Action Models
Fig. 1 ·

Fig. 1 : Policy routing for robot manipulation. We introduce SWAP, an offline reinforcement learning framework for online policy routing. Given candidate robot policies, SWAP dynamically selects the best policy online for an instruction and current observation during task execution.

arXiv

Interpretation

Introduces SWAP, a stepwise policy routing framework that treats choosing which policy to run as an offline reinforcement learning problem over a discrete action space of policy identities, learning a routing critic that estimates each candidate's long-horizon utility from the current observation. Prior VLA deployment assumes one policy for a whole episode, or selects once at episode start; SWAP re-evaluates after every action chunk and can switch policies within a single trajectory, while treating candidate policies as fixed black boxes so new policies can be added without retraining any expert. On three real DROID tasks, SWAP averages 96.7% success versus 63.3% for the best single policy and 66.7% for random routing; evaluation uses 10 held-out trials per task, with training data of 10 successful and 10 failed rollouts per policy per task, 180 trajectories total.

The routing critic learns structured, task-phase-dependent policy preferences rather than random composition: one policy is preferred during early free-space motion, another during precise grasping, and after a dropped object the critic re-selects a policy to recover. This decomposes the gain into switching versus learned switching: random routing beats the best fixed policy by only 3.4 points on average, while SWAP adds a further 30 points over random routing. Q-value analysis reports the weaker standalone policy has a lower value 36% of the time while another policy is assigned the higher value 64% of the time; SWAP changes the selected policy an average of 8.7 times per rollout, each switch corresponding to an average Q-value improvement of +0.117; by normalized trajectory thirds, one policy accounts for 77.4% of selections in the middle third and 63.7% in the final third.

Routing transfers under distribution shift: with unseen distractor objects added to two DROID tasks and no critic retraining, SWAP reaches 90% and 80% success versus 40% and 30% for the best individual policies. This indicates the critic itself generalizes to novel scenes, with the limiting factor being whether candidate strengths remain stable across the shift rather than the routing mechanism failing. On LIBERO-Plus, five heterogeneous policies invert in ranking, and SWAP achieves 62.65% while random routing collapses to 2.7%; removing the policy that is strongest in-distribution but degraded out-of-distribution raises SWAP to 70.65%.

When the candidate set contains one strong and one substantially weaker policy, SWAP identifies and prioritizes the stronger behavior, matching or exceeding that strong policy. This shows routing is not dragged down by low-quality candidates, and it holds when SWAP is trained on only 8 of 10 LIBERO-Spatial tasks and evaluated on two unseen tasks. On LIBERO-Plus, SWAP reaches 89% success, closely matching the stronger standalone policy, while random routing degrades to 79%.

Perspective

The work targets tabletop manipulation settings where multiple pretrained VLA policies exist and offline trajectories with both successes and failures can be collected, and it applies to short-horizon tasks with a small candidate set (three policies in the real experiments, up to five in simulation). Routing happens at action-chunk granularity, aligned with the underlying VLAs' inference cadence, so no additional policy queries are introduced, making it suitable for policy libraries that already emit chunked actions. For engineering teams wanting to compose heterogeneous policies without retraining experts, or to sustain success under distribution shift, the framework offers a directly reusable pipeline: collect per-policy rollouts offline, fit a double-Q critic with IQL, and at deployment select the highest-valued policy.

How learned routing scales to dozens or hundreds of specialized policies remains open, especially under increasingly out-of-distribution scenarios; evaluation is restricted to short-horizon tabletop manipulation, and on longer-horizon tasks the sparse terminal reward becomes a weaker training signal for the routing critic, likely requiring intermediate subtask completion signals; as the candidate set grows, querying multiple heterogeneous policies introduces systems challenges that may make on-device deployment intractable; and the formulation optimizes only task return without accounting for inference latency or monetary cost. In addition, this is a full-text parse, but some figures and tables appear as images, so the specific values in Tables II and III should be checked against the original figures.

Sources