Skip to main content
Back to timeline
arXivSource publication:

BIRD Unifies Iterative Reward Design Evaluation: Simple Design Choices Are Most Robust, and ERA-U/ERA-S Rank First and Second Under Matched Budgets

Related research and updates

Synopsis

The authors introduce BIRD, which expresses 14 published iterative reward design (IRD) methods as configurations of one loop, compares 6 published methods and ablates 30 design choices across MuJoCo, Meta-World, Assistax and HumanoidBench under matched feedback conditions and policy-training budgets, finding that most choices (including the authors' own proposals) have small or inconclusive effects while a few simple choices (such as greedily retaining checkpoints and summarizing feedback) help consistently; a small hill-climb from RDA yields ERA-U and ERA-S, which take first and second place in the pooled Bradley–Terry ranking, though a small human study shows high benchmark scores do not imply natural behavior.

Source-provided article image: A Bird's-Eye View of Iterative Reward Design
Figure 1 ·

Figure 1: Most iterative reward design (IRD) methods follow this basic loop. BIRD uses this shared structure to represent methods in a common format, making them easier to compare.

arXiv

Interpretation

BIRD represents IRD methods as configurations of one loop that reuse implementations of shared design choices, enabling direct comparison, ablation and recombination under matched feedback conditions and policy-training budgets. Previously, differences in implementation details, feedback assumptions, model backbones and policy-training procedures made fair comparison or isolation of individual design choices difficult; BIRD uses shared configurations to prevent implementation differences (such as tie-breaking logic) from becoming confounders. The framework implements 14 published methods and provides evaluation procedures for MuJoCo, Meta-World MT10, Assistax and HumanoidBench; the authors state implementations were verified against original papers and reference source code, though not all papers open-source their code and some model backbones have been deprecated, making complete fidelity difficult.

Comparing 6 published methods under matched conditions, REvolve performs strongest overall (falling behind RDA only on HumanoidBench) and no other method consistently outperforms another with non-overlapping CIs; ablating 30 design choices shows most choices (including the authors' own proposals) have small or inconclusive effects, while a few simple choices help consistently. Prior benchmarks and ablations typically tested design choices within particular algorithms and environments, leaving unclear whether they transfer to other algorithms, environments or combinations; this work evaluates them systematically across tasks and algorithmic stacks, suggesting intuition-driven design choices often fail to transfer. The ablation covers nine tasks (six MT10 tasks and three Gym MuJoCo tasks) with five seeds per experimental cell, expresses effects as Glass's Δ and pools them with 95% BCa intervals over tasks; the authors note that exhaustively evaluating BIRD's more than 300 configurable choices is computationally infeasible, so the 30 choices were selected manually.

A small hill-climb from RDA on four unsaturated MT10 tasks yields ERA-U (unsupervised) and ERA-S (supervised), which take first and second place in the pooled Bradley–Terry ranking, with the ranking holding when development tasks are excluded. ERA-U adds five choices to RDA (including the peak own-reward checkpoint and design thought), while ERA-S flips only the ranking signal from the VLM completion rating to the reference reward with no further tuning; the authors present them not as methods to adopt but as evidence that simple choices plus a small search go a long way. ERA-S outperforms the strongest supervised baseline REvolve and ERA-U outperforms its unsupervised parent RDA, both with non-overlapping CIs; ERA-U never sees the benchmark signal yet ranks above every supervised method except REvolve; the climb used 24 configurations, most run at a reduced budget of 24–30 policy trainings.

A small human study (two authors, 257 clips, 22 tasks) shows high benchmark scores do not imply natural behavior: on Assistax and HumanoidBench none of the evaluated methods (including ERA-U and ERA-S) achieve consistent success and naturalness, and HumanoidBench's walk and crawl reference rewards can be maximized without walking or crawling as a human would. Prior work mostly reported benchmark scores; here completion and naturalness are rated separately, showing naturalness ratings track the reference reward less well than completion ratings and that the VLM judge's deviation nearly doubles on high-scoring policies. Of 67 learned policies, 9 exceed 75% reference reward and raters judged all 9 unnatural; the authors rule out RL merely finding strange policies using two scripted policies (hopping facing backward or rocking in place) that reach comparable reference reward; 70 of 73 HumanoidBench naturalness ratings gave a score of 0, all with high confidence.

Perspective

BIRD targets researchers and engineers who need to compare, ablate or recombine LLM-assisted iterative reward design methods, and applies to robotic-control benchmarks (MuJoCo, Meta-World MT10, Assistax, HumanoidBench) under matched feedback conditions and policy-training budgets. It lets researchers reproduce published methods, modify individual design choices, or combine choices from different methods within one framework, and prototype new components; ERA-U and ERA-S show that a small hill-climb from an existing method can yield top-ranked recipes, and the authors expect further search to surpass them substantially. The human rating portion is a preliminary alignment assessment, indicating that behavior naturalness should be examined alongside benchmark scores.

Human ratings come from two authors and may not represent broader human preferences; only a subset of IRD methods is evaluated, excluding specialized approaches such as LIMEN and Auto-MC; computational efficiency and ways to reduce IRD cost are not studied; ERA-U and ERA-S come from a deliberately small search and are not attempts at state-of-the-art methods; the generator and judge models were not systematically varied, so dependence on model choice is untested, and task information may have leaked into the VLMs since the benchmarks are public; evaluation covers robotic control only. In addition, the VLM judge's deviation nearly doubles on high-scoring policies, meaning the ranking signal for unsupervised methods weakens as policies improve, which could cap their performance; how to encode naturalness in the objective or feedback remains an open problem.

Sources