Skip to main content
Back to timeline
arXivSource publication:

DexPolicy makes exploration noise an explicit function of training steps, lifting deterministic dexterous-manipulation success from 49.4% to 68.1% (FPO) in simulation and on hardware

Synopsis

The work introduces DexPolicy, which in trajectory-guided dexterous manipulation reinforcement learning such as ViViDex makes the action-noise scale an explicit annealing function of training steps while holding loss, architecture, reward, and optimizer settings fixed; across PPO, critic-free GRPO continuation, and a flow-parameterized PPO variant (FPO), mean deterministic Target success over five YCB objects and three training seeds rises from 32.0% to 35.7%, 14.1% to 45.4%, and 49.4% to 68.1%, and on a RealMan RM75 arm with an Inspire/RH56 hand, 360 trials raise mean Target success from 8.3% to 43.3%, 10.0% to 63.3%, and 25.0% to 85.0%.

AI-generated editorial illustration: DexPolicy: Scheduled Exploration for Trajectory-Guided Dexterous Manipulation

Interpretation

Making the exploration scale an explicit function of training steps rather than something learned by the update raises final deterministic Target success in all three policy-optimization settings. Previously the ViViDex PPO baseline learned a Gaussian action standard deviation, and three mustard baseline runs ended at a logged scale of 0.185–0.186 after 5M steps from a nominal 0.20 initialization; this work freezes that parameter and anneals it by step count while keeping the loss form and optimizer settings unchanged. Simulation covers five YCB objects and three training seeds, with PPO and FPO at roughly 5M steps and GRPO as matched checkpoint continuations; each checkpoint receives 100 stochastic and 100 deterministic episodes, totaling 96 checkpoints, 192 evaluations, and 19,200 episodes, reported as means and sample standard deviations without significance claims.

Deterministic execution on hardware shows the same gains, indicating the improvement is not from sampling less Gaussian noise at deployment. All conditions execute policy means, so the hardware differences trace to training-time noise scheduling rather than evaluation-time noise magnitude. On a RealMan RM75 arm with an Inspire/RH56 hand, three objects, 18 object–method conditions, one model and 20 trials each for 360 trials total; PPO rises from 8.3% to 43.3%, GRPO from 10.0% to 63.3%, and FPO from 25.0% to 85.0%, with lift success rising from 38.3% to 58.3%, 46.7% to 95.0%, and 45.0% to 93.3%.

PPO component screening attributes the gain to noise control rather than the tested update contraction, and the selected schedule beats linear decay with the same endpoints. In the single-seed mustard screen, noise scheduling alone raises reward AUC from 19.96 to 27.36 and improves contact and lift, update contraction alone lowers reward AUC to 12.35 and lift AUC to 0.00293 m, and combining both fails to recover the baseline; the shape gate shows uniform linear decay and continuous front-loaded decay both fail, removing the 2M jump passes, and removing the final drop fails with 78.2% lift retention. The component screen is single-seed; the linear control uses three independent training seeds and 100 deterministic episodes per model, giving mustard 75/95/81%, mug 33/26/30%, and banana 1/2/1%, with seeds not paired to the benchmark seeds and the shape result specific to PPO.

Training return, deterministic Target success, and tolerance to execution noise dissociate, so schedules should be judged by terminal task success under the intended execution conditions. FPO on sugar trails before overtaking, PPO's sugar return rises while final Target success stays at zero, and GRPO's mustard return rises despite a mean Target reversal across replication seeds; GRPO's five-object Target at common noise 0.03 contrasts with a reversal at 0.10, and native noise gives 7.0% versus 43.5%. The GRPO continuation study comprises 30 models (five objects, two conditions, three seeds); frozen models receive 200 deterministic and 200 native-stochastic episodes plus 100 at each common standard deviation of 0.03/0.10, totaling 120 evaluations and 18,000 episodes, with per-episode initial-pose equality not recorded.

Perspective

The result applies to the state-policy stage of the ViViDex-style setting, to per-object trained policies, and to deployment that executes policy means deterministically; it is most directly usable by robot-learning researchers and engineering teams seeking higher terminal task success without changing the loss or optimizer. The scheduling rule itself is reused across objects, code and a website are released, and the hardware platform is a RealMan RM75 arm with an Inspire/RH56 hand trained in an embodiment-matched simulator.

Three training seeds limit stability estimates, and hardware uses one model per condition on a single platform, so repeated trials cannot establish training-seed robustness or cross-platform transfer; the mean mustard reversal across GRPO replication seeds limits repeatability; continuations establish neither from-scratch efficiency nor budget-matched cross-method performance; resetting, freezing, and annealing remain unseparated, and the shared credit-assignment rule is not separately ablated; native stochastic evaluations mix policy quality with execution noise, and matched-noise diagnostics cover only GRPO at two scales; three of eight PPO/FPO objects remain effectively unsolved and extended training covers only mustard; ablations cover one optimizer contraction and one fixed FPO scale, and a tuned fixed low-noise PPO baseline is absent. In addition, some table values appear as blanks in the loaded text, so specific cell numbers should be checked against the original.

Sources