Skip to main content
Back to timeline
arXivSource publication:

BVER's bidirectional Voronoi-biased curriculum reaches 95% success on a 0.4 m box climb in roughly 65% fewer iterations than the best reference-free curriculum

Related research and updates

Synopsis

The authors propose BVER, a bidirectional curriculum inspired by bidirectional RRT planning that grows start states outward from the goal and goals outward from the initial state distribution, biases both toward unexplored task space with Voronoi bias, and steers them toward each other, training one goal-conditioned policy on both; on point-mass mazes, quadrupedal box climbing, and robot-arm ring-on-peg transfer it learns faster than all compared reference-free curricula, reaches 95% success on a 0.4 m box in roughly 65% fewer iterations, is the only one of them to learn a 0.7 m box, and approaches the sample efficiency of reference-based curricula on the 0.4 m box and ring-on-peg transfer without a demonstration.

Source-provided article image: Bidirectional Voronoi-biased Exploration Curriculum for Reinforcement Learning
Figure 2 ·

Figure 2: BVER’s expansion mechanism. Intermediate starts s i s_{i} (blue) grow from the target goal g ⋆ g^{\star} and goals g i g_{i} (green) from the initial distribution ρ 0 \rho_{0} . Random walks start from the states nearest to uniform task-space samples, favoring states whose Voronoi cell (shaded) reaches into unexplored space, or nearest to the other direction’s solved states to connect the two (dashed).

arXiv

Interpretation

BVER expands the curriculum from both ends at once: it grows start states outward from the goal and goals outward from the initial state distribution, applies Voronoi bias to both so they point toward unexplored task space, and steers the two ends toward each other, training a single goal-conditioned policy on both. Prior automatic start-state and goal curricula typically expand from one side only, so that side must cover the full distance to the target; BVER borrows bidirectional expansion from bidirectional RRT planning and brings it into reference-free curriculum learning. The abstract states the method's composition and motivation and supports it with an ablation: expanding from both ends outperforms either direction alone.

On point-mass mazes, quadrupedal box climbing, and robot-arm ring-on-peg transfer, BVER learns faster than all compared reference-free curricula. Relative to reference-free curriculum baselines, BVER achieves faster learning across multiple task families rather than in a single environment. The abstract reports comparisons across three task families without giving per-baseline numeric values.

On box climbing, BVER reaches 95% success on a 0.4 m box in roughly 65% fewer iterations, is the only compared method to learn a 0.7 m box, and yields a policy robust to start, goal, and yaw variation. Compared with the best reference-free curriculum baseline, BVER gains both iteration efficiency and reachable box height, plus robustness to start, goal, and yaw perturbations. The abstract reports the 95% success rate, the roughly 65% iteration reduction, the unique 0.7 m box success, and the robustness description, all at abstract level.

Without a demonstration, BVER approaches the sample efficiency of reference-based curricula on the 0.4 m box and on ring-on-peg transfer. This indicates that a bidirectional automatic curriculum can narrow the sample-efficiency gap to reference-based methods without reference motions or demonstrations. The abstract describes the comparison as approaching, without giving sample counts or curve values.

Perspective

The work targets reference-free, sparse-reward, long-horizon goal-conditioned reinforcement learning, aimed at researchers and engineering teams who want to avoid demonstration collection and task-specific reward engineering; its validation settings are point-mass mazes, quadrupedal box climbing, and robot-arm ring-on-peg transfer, with box-climbing results reported as 95% success on a 0.4 m box, roughly 65% fewer iterations, and unique success on a 0.7 m box. Methodologically, the combination of Voronoi bias and bidirectional expansion offers a reusable way to port two-sided planning ideas into curriculum generation.

The abstract does not give per-baseline numeric values, absolute iteration counts, the number of random seeds or statistical variation, nor the concrete implementation of the Voronoi bias, hyperparameter sensitivity, or real-robot performance; the claim that BVER is the only method to learn the 0.7 m box is limited to the compared reference-free curricula. Confirming these points requires the full text and figures, so this summary rests on abstract-level information.

Sources