Public articles linked to the same research event.
arXiv The authors introduce BIRD, which expresses 14 published iterative reward design (IRD) methods as configurations of one loop, compares 6 published methods and ablates 30 design choices across MuJoCo, Meta-World, Assistax and HumanoidBench under matched feedback conditions and policy-training budgets, finding that most choices (including the authors' own proposals) have small or inconclusive effects while a few simple choices (such as greedily retaining checkpoints and summarizing feedback) help consistently; a small hill-climb from RDA yields ERA-U and ERA-S, which take first and second place in the pooled Bradley–Terry ranking, though a small human study shows high benchmark scores do not imply natural behavior.
The authors introduce BIRD, which expresses 14 published iterative reward design (IRD) methods as configurations of one loop, compares 6 published methods and ablates 30 design choices across MuJoCo, Meta-World, Assistax and HumanoidBench under matched feedback conditions and policy-training budgets, finding that most choices (including the authors' own proposals) have small or inconclusive effects while a few simple choices (such as greedily retaining checkpoints and summarizing feedback) help consistently; a small hill-climb from RDA yields ERA-U and ERA-S, which take first and second place in the pooled Bradley–Terry ranking, though a small human study shows high benchmark scores do not imply natural behavior.
The authors introduce BIRD, which expresses 14 published iterative reward design (IRD) methods as configurations of one loop, compares 6 published methods and ablates 30 design choices across MuJoCo, Meta-World, Assistax and HumanoidBench under matched feedback conditions and policy-training budgets, finding that most choices (including the authors' own proposals) have small or inconclusive effects while a few simple choices (such as greedily retaining checkpoints and summarizing feedback) help consistently; a small hill-climb from RDA yields ERA-U and ERA-S, which take first and second place in the pooled Bradley–Terry ranking, though a small human study shows high benchmark scores do not imply natural behavior.
The authors introduce BIRD, which expresses 14 published iterative reward design (IRD) methods as configurations of one loop, compares 6 published methods and ablates 30 design choices across MuJoCo, Meta-World, Assistax and HumanoidBench under matched feedback conditions and policy-training budgets, finding that most choices (including the authors' own proposals) have small or inconclusive effects while a few simple choices (such as greedily retaining checkpoints and summarizing feedback) help consistently; a small hill-climb from RDA yields ERA-U and ERA-S, which take first and second place in the pooled Bradley–Terry ranking, though a small human study shows high benchmark scores do not imply natural behavior.