A Dirichlet-process-mixture skill prior lets long-horizon robot manipulation exceed 0.8 success within 1.5M steps, while SAC stays below 0.1 after 5M
Synopsis
The work proposes a Bayesian non-parametric skill prior built on a Dirichlet Process Mixture that models temporally extended skills in a structured latent space, enabling adaptive skill discovery without predefining the number of components; integrated into a hierarchical RL framework, the learned prior guides high-level skill selection while a pretrained decoder generates temporally abstracted actions, and experiments on Franka Kitchen, LIBERO-Long, Meta-World, and a real robot show consistent gains in long-horizon manipulation, achieving over 0.8 success rate within 1.5M steps whereas SAC fails to converge even after 5M steps (below 0.1), with an average improvement of 21.8% over a single-Gaussian prior baseline.
Fig. 1: HELIOS Framework overview. The training process is divided into two phases. In Phase I, a VAE with GRU modules is used to pre-train a skill representation model from a dataset of action trajectories. The model leverages a DPM to capture the non-parametric nature of skill priors, aiding in learning precise action patterns and subsequent effective task representations. In Phase II, this pre-trained skill decoder and prior are deployed within a RL framework to address long-horizon manipulation tasks. Here, the upstream inference model uses soft actor-critic structure to learn specific task reasoning, ensuring the successful execution of complex, extended long-horizon tasks.
arXivInterpretation
A Bayesian non-parametric skill prior is built with a Dirichlet Process Mixture to represent temporally extended skills in a structured latent space, so the number of skill components is determined adaptively rather than predefined. Existing skill-based approaches often assume a fixed parametric prior such as a single Gaussian, which limits their ability to capture the multi-modal skill structures complex tasks require; replacing that fixed prior with a non-parametric mixture removes the need to preset the component count. The abstract states the prior is integrated into a hierarchical RL framework, where the learned prior guides high-level skill selection and a pretrained decoder generates temporally abstracted actions, improving exploration efficiency under sparse rewards.
On sparse-reward long-horizon manipulation, the method shows consistent gains across Franka Kitchen, LIBERO-Long, Meta-World, and a real robot, reaching over 0.8 success rate within 1.5M steps. As a comparison, SAC fails to converge even after 5M steps with success below 0.1, indicating a clear advantage in settings with delayed feedback and inefficient exploration. The abstract reports consistent gains across four evaluation settings including a real robot, with concrete numbers for success rate and training steps.
Relative to a single-Gaussian prior baseline, the model yields an average improvement of 21.8%. This comparison directly targets the limitation of fixed parametric priors and quantifies the benefit of a non-parametric multi-modal skill prior over a single-Gaussian prior. The abstract reports an average improvement of 21.8%, which is an average across tasks; no per-task breakdown is given in the abstract.
Perspective
The result targets long-horizon robot manipulation under sparse rewards, in settings that use hierarchical RL and can rely on a pretrained decoder; it benefits researchers and engineering teams that need to model multi-modal skill structure. The evidence described in the abstract covers Franka Kitchen, LIBERO-Long, Meta-World, and a real robot, indicating the method is examined on both simulated benchmarks and a real platform, though the abstract does not describe the concrete form or scale of the real-robot task.
The abstract does not give per-task success rates, the task setup or trial scale of the real-robot experiments, or how the Dirichlet Process Mixture adapts its component count during training; the 21.8% figure is an average across tasks whose spread and variance are unknown. The abstract also does not report comparisons against a wider set of skill-prior methods, so the relative advantage of this non-parametric prior over other multi-modal priors remains an open question.
