Author treats all four layers of a humanoid learning pipeline as proxies, reporting a curriculum gate that advanced at its rate limit while the gated skill was absent and a synchronised push test that ranked a 0.5 m/s push as more dangerous than a 2.0 m/s one
Related research and updatesSynopsis
The author argues that the reward, curriculum gate, evaluation statistic, and reference motion in a legged-robot reinforcement-learning pipeline are all proxies in the same formal sense, each with a characteristic divergence mechanism and a reformulation that closes it, illustrated by measurements from one continuous lineage of a PPO policy for a simulated 1.91 m humanoid grown over four stages and 13,500 iterations, among which a curriculum gate built on averaged error advanced at its rate limit on every check while the skill it gated was absent, and a batched push test whose synchronised resets aliased the gait phase ranked a 0.5 m/s push as more dangerous than a 2.0 m/s one.
Fig. 1: A curriculum gated on averaged tracking error. Top: the reference amplitude advanced at its rate limit at every check. Bottom: the gated statistic stayed far below its 0.25 rad threshold, although the principal joint’s peak error was ≈ \approx 1 rad.
arXivInterpretation
The paper treats the reward, curriculum gate, evaluation statistic, and reference motion as proxies in the same formal sense, locating specification failure beyond the reward alone. The traditional view treats only the reward as optimised against and places reward hacking in the reward itself; this work extends proxy divergence to the curriculum gate, evaluation statistic, and reference motion. The argument is primarily formal derivation, illustrated by measurements from one continuous lineage of a PPO policy for a simulated 1.91 m humanoid grown over four stages and 13,500 iterations.
Each layer has a characteristic divergence mechanism and admits a reformulation that closes it. For every layer the paper states the traditional formulation, derives the divergence condition, and gives an alternative: first-order (L1) costs where quadratic kernels are flat, peak and outcome statistics where curriculum gates average, gate reachability and information checks, deterministic and phase-desynchronised evaluation, curriculum state treated as part of the model, feasibility-first reference design with residual feed-forward, and function-preserving input widening that lets one policy grow instead of being retrained. The alternatives are presented as derivations and illustrated by measurements from the same policy lineage.
A curriculum gate built on averaged error advanced at its rate limit on every check while the skill it gated was absent. This gives a concrete observable form of curriculum-gate divergence as a proxy, rather than only a theoretical possibility. From measurements of the PPO policy lineage for the simulated 1.91 m humanoid across four stages and 13,500 iterations; the paper reports the gate advanced at its rate limit on every check.
In a batched push test, synchronised resets aliased the gait phase, so a 0.5 m/s push was ranked as more dangerous than a 2.0 m/s one. This shows an evaluation statistic acting as a proxy can produce a directionally wrong ordering, not merely a numerical offset. From the batched push test measurements in the same policy lineage; the paper reports this ranking reversal.
Perspective
The work addresses researchers and engineers assembling legged-robot reinforcement-learning pipelines, in staged training and evaluation settings where the reward, curriculum gate, evaluation statistic, and reference motion each play a proxy role. Its reformulations are organised layer by layer, so they can be substituted into an existing pipeline: first-order (L1) costs where quadratic kernels are flat, peak and outcome statistics where curriculum gates average, gate reachability and information checks, deterministic and phase-desynchronised evaluation, curriculum state treated as part of the model, feasibility-first reference design with residual feed-forward, and function-preserving input widening that lets one policy grow. The illustrative measurements come from one continuous lineage of a PPO policy for a simulated 1.91 m humanoid, grown over four stages and 13,500 iterations on a single laptop GPU.
A careful reader would still watch how these divergence mechanisms behave on other robot morphologies, other algorithms, or real hardware, since the reported measurements come from one PPO policy lineage for a simulated 1.91 m humanoid over four stages and 13,500 iterations. The reproducibility conditions and statistical stability of the two observations, the curriculum gate advancing at its rate limit while the gated skill was absent and the 0.5 m/s push being ranked as more dangerous than the 2.0 m/s one, remain open to testing across more settings. In addition, this reading covers only the abstract, without the body, figures, or appendices, so the full derivations of each layer's divergence condition, the implementation details of the alternatives, and the complete measurement set cannot be summarised here; these are open questions to check against the original.
