Public articles linked to the same research event.
arXiv The author argues that the reward, curriculum gate, evaluation statistic, and reference motion in a legged-robot reinforcement-learning pipeline are all proxies in the same formal sense, each with a characteristic divergence mechanism and a reformulation that closes it, illustrated by measurements from one continuous lineage of a PPO policy for a simulated 1.91 m humanoid grown over four stages and 13,500 iterations, among which a curriculum gate built on averaged error advanced at its rate limit on every check while the skill it gated was absent, and a batched push test whose synchronised resets aliased the gait phase ranked a 0.5 m/s push as more dangerous than a 2.0 m/s one.
The author argues that the reward, curriculum gate, evaluation statistic, and reference motion in a legged-robot reinforcement-learning pipeline are all proxies in the same formal sense, each with a characteristic divergence mechanism and a reformulation that closes it, illustrated by measurements from one continuous lineage of a PPO policy for a simulated 1.91 m humanoid grown over four stages and 13,500 iterations, among which a curriculum gate built on averaged error advanced at its rate limit on every check while the skill it gated was absent, and a batched push test whose synchronised resets aliased the gait phase ranked a 0.5 m/s push as more dangerous than a 2.0 m/s one.
The author argues that the reward, curriculum gate, evaluation statistic, and reference motion in a legged-robot reinforcement-learning pipeline are all proxies in the same formal sense, each with a characteristic divergence mechanism and a reformulation that closes it, illustrated by measurements from one continuous lineage of a PPO policy for a simulated 1.91 m humanoid grown over four stages and 13,500 iterations, among which a curriculum gate built on averaged error advanced at its rate limit on every check while the skill it gated was absent, and a batched push test whose synchronised resets aliased the gait phase ranked a 0.5 m/s push as more dangerous than a 2.0 m/s one.
The author argues that the reward, curriculum gate, evaluation statistic, and reference motion in a legged-robot reinforcement-learning pipeline are all proxies in the same formal sense, each with a characteristic divergence mechanism and a reformulation that closes it, illustrated by measurements from one continuous lineage of a PPO policy for a simulated 1.91 m humanoid grown over four stages and 13,500 iterations, among which a curriculum gate built on averaged error advanced at its rate limit on every check while the skill it gated was absent, and a batched push test whose synchronised resets aliased the gait phase ranked a 0.5 m/s push as more dangerous than a 2.0 m/s one.