ReCo cuts simulated end-effector tracking error by about 28% with response-consistent locomotion and policy-aware MPC
Synopsis
ReCo couples a response-consistent reinforcement-learning locomotion policy with policy-aware model predictive control (MPC): response shaping trains the policy to respond to commands consistently and repeatably across randomized dynamics, and an identified closed-loop response model then lets MPC jointly plan locomotion commands and arm motion, reducing position and orientation root-mean-square error (RMSE) on the simulation benchmark by 28.7% and 27.4% relative to the best baseline for each metric, with real-world experiments demonstrating onboard continuous legged manipulation and coordinated base and arm motion.
Fig. 1: Demonstration of continuous legged manipulation tasks in the real world. ReCo jointly plans locomotion commands and arm motions to track task-space end-effector trajectories for object organization, cabinet-door opening, and pick-and-place.
arXivInterpretation
It presents ReCo, a framework coupling response-consistent legged locomotion with policy-aware MPC for continuous end-effector tracking while the base keeps walking in legged manipulation. Relative to existing RL-plus-MPC combinations, it explicitly addresses the fact that a learned policy's command response varies with gait phase, contact, and payload, so that MPC can predict and compensate for base motion. At the abstract level it reports the framework components and simulation benchmark results, without ablation, sample-size, or statistical-test details.
Response shaping trains the policy to respond to commands consistently and repeatably across randomized dynamics, after which a closed-loop response model is identified for MPC. It makes response consistency a training objective and derives a closed-loop response model usable by MPC, rather than treating the policy as a fixed black box. The abstract states the method steps but gives no randomization ranges, identification data volume, or fit accuracy.
MPC jointly plans locomotion commands and arm motion to compensate for tracking errors. It brings base locomotion commands and arm motion into a single plan, rather than relying on the arm to compensate passively. Method description at the abstract level, without planning horizon, frequency, or constraint form.
On the simulation benchmark it reduces position and orientation RMSE by 28.7% and 27.4% relative to the best baseline for each metric, and real-world experiments demonstrate onboard continuous legged manipulation with coordinated base and arm motion. It provides quantified improvement over the best baseline in simulation and validates onboard continuous manipulation on a real platform. The abstract reports percentage improvements and the existence of real-world experiments, without a baseline list, trial counts, or quantitative real-world metrics.
Perspective
The work targets legged manipulation settings that require accurate end-effector tracking while the base keeps walking, where the policy response varies with gait phase, contact, and payload and MPC needs predictable base motion. Its contribution is a reusable combination: response shaping first yields consistent, repeatable command responses, and an identified closed-loop response model then supports MPC in jointly planning locomotion commands and arm motion. For researchers and engineers working on legged robots, mobile manipulation, and hybrid learning-control and MPC architectures, this framework offers a reference path for explicitly bringing policy response characteristics into the planner; the simulation RMSE reductions and the real onboard demonstration indicate the path runs in both simulation and on a real platform.
The abstract does not state the range of randomized dynamics, the identification data and accuracy of the closed-loop response model, the MPC planning horizon and frequency, the baseline set, or trial counts, nor does it give quantitative real-world tracking metrics. The specific baselines and evaluation conditions behind the 28.7% and 27.4% RMSE reductions, and the accuracy of coordinated motion on the real platform, therefore still need confirmation from the full text. How response consistency is implemented and what it costs in training, and how far the method applies under payload or terrain variation, are questions a reader can continue to watch.
