Skip to main content
Back to timeline
arXivSource publication:

Learning Holistic Whole-Body Loco-Manipulation with a Bipedal Mobile Manipulator

Synopsis

This work presents a unified reinforcement-learning whole-body controller that takes a single 6-DoF end-effector target as the sole task-level command and directly outputs coordinated actions for a bipedal base and a six-joint arm across 14 joints, reaching 88.30% task success in simulation (2.85 cm mean position error, 5.38 cm P95), and on a real bipedal platform reusing the same controller across VR teleoperation, a learned diffusion policy, and scripted trajectories while extending vertical reach from roughly 38-163 cm for a floating-base plus inverse-kinematics baseline to roughly 3-191 cm.

Source-provided article image: Learning Holistic Whole-Body Loco-Manipulation with a Bipedal Mobile Manipulator
Fig. 1 ·

Fig. 1: Overview of the proposed controller. Given a 6-DoF end-effector target and proprioceptive observations, the policy coordinates whole-body actions through reward-gated training and temporal context estimation.

arXiv

Interpretation

It formulates and realizes an end-effector-driven whole-body control interface for a bipedal manipulator, where the controller itself decides when arm motion alone suffices and when postural adaptation or stepping is needed. Unlike interfaces that require explicit root-velocity, torso, or gait commands such as ULC and CEER, or that handle locomotion with a separate policy as in Portela et al., this work unifies mobility and manipulation under a single 6-DoF end-effector target, with one policy jointly controlling all 14 arm and leg joints rather than decomposing upper- and lower-body control. The method section provides a complete training pipeline and a deployable input table (58-dimensional observation, 9-dimensional end-effector command, 67-dimensional context), and the same controller is reused without modification across three command sources on the real robot.

It introduces biped-aware reward gating: a scheduled SE(3)-distance reference produces a continuous locomotion-manipulation phase coefficient, combined with phase-specific safety scores, a lower-bound offset c_s = 0.4, and a best-so-far progress reward. Building on the RFM of Jiang et al., it replaces the cumulative-error term with progress relative to the best errors achieved so far and adds a lower bound to the safety score so the learning signal is retained during recovery. In simulation ablations, compared with a matched additive reward using the same terms and coefficients, success rises from 82.73% to 88.30%, mean position error falls from 3.23 cm to 2.85 cm, position P95 falls from 14.35 cm to 5.38 cm, and action variation falls from 0.218 to 0.165; the additive reward does achieve lower orientation error, indicating a trade-off between individual tracking terms and task-level whole-body performance.

It proposes a Transformer-GRU temporal context estimator that combines windowed attention with recurrent memory plus next-observation prediction and velocity supervision to infer dynamics-relevant information from proprioceptive history. Relative to no-latent, DreamWaQ/CENet, GRU-only, and Transformer-only variants, the full estimator attains the highest non-oracle success rate. Ablations show that removing temporal context drops success to 69.87% and raises action variation to 0.236; the full estimator reaches 88.30%, 3.87 percentage points above Transformer-only, with a privileged oracle at 94.53% as an approximate upper bound; t-SNE and LDA visualizations show the latent space organized by locomotion, approach, and precision-tracking phases with continuous transitions between adjacent phases.

It integrates and deploys the controller on a low-cost bipedal single-arm platform, validating the unified interface and a larger manipulation workspace. Compared with a floating-base plus inverse-kinematics baseline, the controller extends vertical reach without explicit base-posture or locomotion commands and maintains smoother whole-body behavior near workspace boundaries. The real platform comprises a LimX TRON 1 biped, an ARX L5 arm, a UMI gripper, a Livox Mid-360 LiDAR, and a Jetson Orin NX, with the policy at 50 Hz and low-level PD tracking at 500 Hz; demonstrations include picking a plush toy from the ground, stepping, and placing it on a shelf, wiping a blackboard, a diffusion policy pushing a cabinet door closed, and scripted up-and-down trajectories.

Perspective

The results target a minimalist bipedal single-arm platform with 14 controlled joints, where the controller takes an instantaneous 6-DoF end-effector target as the sole external task command, is trained in Isaac Lab simulation, and is transferred to the real robot. Its value lies in providing a common execution interface for VR teleoperation, learned policies, and scripted trajectories, and in extending vertical reach. It applies to manipulation tasks that require reaching objects at different heights and locations and that permit squatting, extension, and stepping in exchange for reachability.

Simulation ablations compare methods under a shared target distribution and domain-randomization setting, while the real-robot portion centers on demonstrations and a reachability comparison rather than large-scale real-world success statistics; the latent visualization shows overlap between adjacent phases, indicating continuous rather than discrete phase categories; the authors state that future work will target more dynamic loco-manipulation behaviors and improved task-space tracking precision for fine-grained manipulation, especially when following learned high-level commands, and those directions remain to be validated.

Sources