Public articles linked to the same research event.
arXiv The work presents a jointly trained two-level reinforcement learning framework in which a high-level SAC policy observes robot pose, 36 raycasts, moving-obstacle states and a receding-horizon local goal and emits a body-velocity command every ten control steps, while a velocity-conditioned low-level SAC gait policy tracks each command through PD joint targets; across 100 evaluation trials per method in randomized PyBullet environments, the method reaches the goal in 98.0% of static and 88.0% of dynamic trials, against at most 78.0% and 68.0% for the SAC+A⋆, SAC+RRT⋆ and SAC+APF hybrids, with path lengths within 4% of the A⋆ reference, and ablations show each observation channel and reward term contributes materially.
The work presents a jointly trained two-level reinforcement learning framework in which a high-level SAC policy observes robot pose, 36 raycasts, moving-obstacle states and a receding-horizon local goal and emits a body-velocity command every ten control steps, while a velocity-conditioned low-level SAC gait policy tracks each command through PD joint targets; across 100 evaluation trials per method in randomized PyBullet environments, the method reaches the goal in 98.0% of static and 88.0% of dynamic trials, against at most 78.0% and 68.0% for the SAC+A⋆, SAC+RRT⋆ and SAC+APF hybrids, with path lengths within 4% of the A⋆ reference, and ablations show each observation channel and reward term contributes materially.
The work presents a jointly trained two-level reinforcement learning framework in which a high-level SAC policy observes robot pose, 36 raycasts, moving-obstacle states and a receding-horizon local goal and emits a body-velocity command every ten control steps, while a velocity-conditioned low-level SAC gait policy tracks each command through PD joint targets; across 100 evaluation trials per method in randomized PyBullet environments, the method reaches the goal in 98.0% of static and 88.0% of dynamic trials, against at most 78.0% and 68.0% for the SAC+A⋆, SAC+RRT⋆ and SAC+APF hybrids, with path lengths within 4% of the A⋆ reference, and ablations show each observation channel and reward term contributes materially.
The work presents a jointly trained two-level reinforcement learning framework in which a high-level SAC policy observes robot pose, 36 raycasts, moving-obstacle states and a receding-horizon local goal and emits a body-velocity command every ten control steps, while a velocity-conditioned low-level SAC gait policy tracks each command through PD joint targets; across 100 evaluation trials per method in randomized PyBullet environments, the method reaches the goal in 98.0% of static and 88.0% of dynamic trials, against at most 78.0% and 68.0% for the SAC+A⋆, SAC+RRT⋆ and SAC+APF hybrids, with path lengths within 4% of the A⋆ reference, and ablations show each observation channel and reward term contributes materially.