Hierarchical RL lets an underactuated biped walk around obstacles: 98.0% static and 88.0% dynamic goal arrival, beating A⋆, RRT⋆ and APF hybrids
Related research and updatesSynopsis
The work presents a jointly trained two-level reinforcement learning framework in which a high-level SAC policy observes robot pose, 36 raycasts, moving-obstacle states and a receding-horizon local goal and emits a body-velocity command every ten control steps, while a velocity-conditioned low-level SAC gait policy tracks each command through PD joint targets; across 100 evaluation trials per method in randomized PyBullet environments, the method reaches the goal in 98.0% of static and 88.0% of dynamic trials, against at most 78.0% and 68.0% for the SAC+A⋆, SAC+RRT⋆ and SAC+APF hybrids, with path lengths within 4% of the A⋆ reference, and ablations show each observation channel and reward term contributes materially.
Fig. 1: Control paradigms for obstacle-aware bipedal locomotion. Flat RL entangles navigation with gait, decoupled planning ignores balance coupling, and the proposed hierarchy learns a velocity-command policy over a learned velocity-conditioned gait.
arXivInterpretation
A jointly trained two-level HRL framework is proposed in which a high-level SAC policy outputs body-velocity commands from pose, 36 raycasts, dynamic-obstacle states and a local goal, and a velocity-conditioned low-level SAC gait policy executes each command for ten control steps. Unlike flat end-to-end RL that maps observations directly to joint targets, and unlike decoupled pipelines that track a geometrically planned path with an independent gait controller, this work separates navigation from gait generation while keeping the feedback between them, shortening the effective navigation horizon tenfold. The paper gives complete definitions of both levels' observations, actions and rewards, and states that the two levels update at 50 Hz and 5 Hz within the same episodes; training curves show the dynamic-obstacle condition converges more slowly because the policy must additionally learn velocity-aware avoidance through the TTC term.
The high-level observation and reward design covers both static layouts and moving obstacles, including a time-to-collision (TTC) term that induces anticipatory avoidance. Compared with reactive methods that rely on instantaneous proximity alone and with global planners that see moving obstacles only at replanning instants, the TTC term penalizes predicted collisions before proximity alone would trigger. Ablations show that removing the dynamic-obstacle velocity channel or obstacle positions severely degrades both regimes, and that removing the TTC or proximity penalty has a comparable effect, confirming anticipatory and proximal safety signals are complementary.
Because the converged low-level gait is task-agnostic, it can be frozen and driven by classical planners over the same command interface, yielding three controlled baselines, SAC+A⋆, SAC+RRT⋆ and SAC+APF, in which only the command generator differs. Prior HRL locomotion studies do not benchmark the learned high-level policy against classical planners driving the identical learned gait, so this design isolates the navigation layer's contribution. With 100 evaluation trials per method per condition, the method attains 98.0% static and 88.0% dynamic success, versus 76.0% to 78.0% static and 68.0% dynamic for the global planners, while APF drops to 55.0% dynamic success with 45.0% collisions; path lengths of 7.92 m static and 7.90 m dynamic stay within 4% of the A⋆ reference.
An ablation over high-level observation channels and reward components quantifies their effect on success, collision and failure rates. The ablation attributes the overall advantage to the coupled navigation-balance structure rather than to any single reward term. Removing the posture terms of Eq. (8) shifts failures toward falls, and removing the shaping clip of Eq. (5) destabilizes training outright (90.0% fall/timeout, static); across variants the dominant failure mode is fall or timeout rather than collision.
Perspective
The result applies to an underactuated biped platform with a known static occupancy map, eight actuated leg joints and no hip or ankle roll joints, evaluated in PyBullet simulation with goals at least 6.0 m away and moving-obstacle trajectories unknown in advance. It makes joint training of a navigation layer and a gait layer a reusable paradigm: the converged velocity-conditioned gait can be frozen and driven by classical planners over the same interface, which isolates navigation intelligence as the comparison target. For a reader, the directly reusable elements are the division of labor between high-level observations (pose, raycasts, dynamic-obstacle states, local goal) and rewards (path, posture, safety, regularization, terminal), the use of a TTC term for anticipatory avoidance, and the use of a posture term to make the navigation layer accountable for gait feasibility.
The study is simulation-only, and sim-to-real transfer would require addressing model mismatch, sensing noise and inter-level latency; the raycasts and dynamic-obstacle states of Eq. (4) are privileged simulator quantities, and how well they stand in for a real LiDAR and object tracker remains to be validated; the planning layer assumes a known static map shared symmetrically by the method and the baselines; because the low-level gait is co-trained with the high-level policy, the frozen gait reused by baselines may track their commands slightly less faithfully; results use a single training seed, so seed variance is unquantified; and the policy actuates only leg joints, leaving upper-body and multi-contact avoidance open. In addition, several equations and table values are missing from the parsed text (for example the discount factor, learning rate, replay buffer size, local-goal lookahead arc length and reward weights), so those specific hyperparameter values cannot be given in this summary.
