Skip to main content
Back to timeline
arXivSource publication:

Modeling Go1 as a cell complex with Hodge message passing yields the highest return on unseen actuator degradations under degradation training

Synopsis

The authors represent the Unitree Go1 as a cell complex whose four limb-level and one body-level rank-2 cells make multi-joint mechanical units explicit, and apply Hodge-based message passing over nodes, edges, and faces; under degradation training, the node-edge-face Hodge actor achieves the highest return on unseen actuator degradations, with higher survival and lower velocity-tracking error.

Source-provided article image: Higher-Order Morphology Priors for Quadruped Reinforcement Learning Under Actuator Degradation
Figure 1 ·

Figure 1 : Go1 higher-order morphology. (a) The twelve actuated joints form the 0-cells; sixteen 1-cells give eight serial couplings, four around the hip ring, and four dashed virtual closures that let each leg bound a face. The five shaded 2-cells are one per limb plus the body cell spanned by the four coplanar hips. (b) The same cells at the true Go1 joint positions, with distant limbs faded. Only the node-edge-face actor propagates through rank-2 cells. The legend maps each mark to its rank: a ring is a 0-cell, a solid line a 1-cell, a dashed line a virtual 1-cell, and a filled patch a 2-cell.

arXiv

Interpretation

A higher-order morphology representation for the Unitree Go1: 12 nodes correspond to the 12 joints, 16 edges connect adjacent joints along the legs and around the hip ring, and four limb cycles plus one body cycle are promoted to 5 rank-2 faces, so limb- and body-level multi-joint units become explicit higher-order objects rather than being implicit in their bounding edges. Prior morphology-aware policies mostly encode morphology as a classical graph with joints as nodes and links as edges, which can represent the edges bounding these structures but not the structures themselves as higher-order objects; this work makes limb- and body-level structure explicit via a cell complex with fixed incidence matrices. The representation is motivated directly by Go1 mechanics: foot placement depends jointly on all three joints of a leg, so lost actuation at one must be compensated by the others, and the four hips are rigidly mounted in a fixed coplanar arrangement forming a configuration-independent hip ring. The stated scale is 12 nodes, 16 edges, and 5 faces.

A node-edge-face Hodge actor: each propagation layer combines same-rank messages, defined by the lower and upper components of the Hodge Laplacian that couple rank-k cells sharing a boundary or a coface, with cross-rank messages aggregated from adjacent ranks through incidence operators, balanced by a learned mixing coefficient and passed through a learned norm gate and RMSNorm before a residual update; after four layers, the same two-layer decoder is applied independently per joint to produce tanh-transformed Gaussian distributions forming a factorized policy. Higher-order graph learning has largely served general signal processing on simplicial or cellular complexes; here rank-2 cells carry persistent mechanical structures of a quadruped and higher-order message passing acts as a policy-level morphology prior evaluated under actuator degradation. All actor variants share the same critic, reward, domain randomization, optimization setup, and training budget, with each policy trained for 400M environment steps; the critic is the standard MuJoCo Playground privileged-observation critic, used for all actor variants to isolate differences among actor architectures.

In a capacity-matched healthy/degradation training-evaluation study, Hodge-F under degradation training achieves the highest return of 22.40, survival 0.94, and lowest velocity RMSE of 0.252 on unseen compound actuator degradations, versus MLP-C at 19.98, 0.88, and 0.324, and Hodge-L at 18.61, 0.83, and 0.325. Ablations indicate the gain is not explained by per-joint factorization alone (MLP-D: return 18.15, velocity RMSE 0.320) nor by a conventional graph (GCN: return 15.35, velocity RMSE 0.339); the advantage of faces over Hodge-L emerges primarily when learned compensation is required rather than under nominal locomotion. Degradation scales the maximum available torque of two actuators at a time, giving 66 unordered pairs split into 44 training and 22 held-out pairs; during training the pair is fixed per episode and sampled as healthy, weak, or dead, and evaluation averages over all held-out pairs and all degradation levels.

Cross-regime evaluation shows conditionality: a policy trained only under the healthy regime gives MLP-C a higher return of 16.89 than Hodge-F at 15.59 under zero-shot degradation, indicating that morphology-aware structure alone is insufficient to guarantee robustness to unseen actuator failures; a policy trained under degradation gives Hodge-F a return of 27.17 and velocity RMSE of 0.220 under nominal evaluation, surpassing MLP-C at 24.85 and 0.306. This suggests that exposure to actuator degradation during training enables the node-edge-face policy to learn a broadly effective locomotion strategy rather than one specialized only to damaged conditions, while bounding the benefit of the higher-order morphology prior to settings where training covers degradation. Cross-regime evaluation swaps the evaluation regimes without retraining; under the healthy regime all actors reach survival 1.00, whereas under the degraded regime survival differences are pronounced.

Perspective

The result concerns velocity-commanded quadruped locomotion in the Unitree Go1 setting of the MuJoCo Playground Go1JoystickFlatTerrain task, where degradation scales the maximum available torque of individual actuators and two actuators are degraded at a time during training. For researchers who want whole-body compensation without adding dedicated fault-handling mechanisms, the reusable elements are the representation that promotes multi-joint units such as limbs and the body to rank-2 cells, and the propagation and decoding structure of the node-edge-face Hodge actor. The benefit applies when the training distribution includes actuator degradation; with healthy-only training, morphology-aware structure alone did not yield robustness to unseen failures.

Open questions a careful reader would still watch: the conclusions rest on a single robot platform and a single simulated task, so applicability to other morphologies, other tasks, or real hardware remains to be tested; the degradation model only scales the maximum available torque and does not cover other failure modes; and cross-regime evaluation shows the benefit depends on whether degradation was seen in training, so the target degradation distribution should be checked against training coverage before deployment. The source was provided as a fast parse, and the contents of Fig. 1 and Fig. 3 are not included in the text, so the cell-complex visualization and learning-curve details cannot be verified here.

Sources