Public articles linked to the same research event.
arXiv The work introduces General Reward Inference and Disentanglement (GRID), a social learning method that uses an information bottleneck to decompose each agent's reward function into a general reward shared across all agents and specific rewards capturing individual preferences, so that training solely on the general reward yields a generalist agent internalizing universal environmental competencies such as safety and basic task proficiency; experiments on a synthetic basis function decomposition, multi-agent Craftax, MuJoCo Gym continuous control, and the Highway-Env driving simulator confirm semantically meaningful reward disentanglement, outperformance of standard learning-from-demonstration baselines, and more efficient and stable specialization.
The work introduces General Reward Inference and Disentanglement (GRID), a social learning method that uses an information bottleneck to decompose each agent's reward function into a general reward shared across all agents and specific rewards capturing individual preferences, so that training solely on the general reward yields a generalist agent internalizing universal environmental competencies such as safety and basic task proficiency; experiments on a synthetic basis function decomposition, multi-agent Craftax, MuJoCo Gym continuous control, and the Highway-Env driving simulator confirm semantically meaningful reward disentanglement, outperformance of standard learning-from-demonstration baselines, and more efficient and stable specialization.
The work introduces General Reward Inference and Disentanglement (GRID), a social learning method that uses an information bottleneck to decompose each agent's reward function into a general reward shared across all agents and specific rewards capturing individual preferences, so that training solely on the general reward yields a generalist agent internalizing universal environmental competencies such as safety and basic task proficiency; experiments on a synthetic basis function decomposition, multi-agent Craftax, MuJoCo Gym continuous control, and the Highway-Env driving simulator confirm semantically meaningful reward disentanglement, outperformance of standard learning-from-demonstration baselines, and more efficient and stable specialization.
The work introduces General Reward Inference and Disentanglement (GRID), a social learning method that uses an information bottleneck to decompose each agent's reward function into a general reward shared across all agents and specific rewards capturing individual preferences, so that training solely on the general reward yields a generalist agent internalizing universal environmental competencies such as safety and basic task proficiency; experiments on a synthetic basis function decomposition, multi-agent Craftax, MuJoCo Gym continuous control, and the Highway-Env driving simulator confirm semantically meaningful reward disentanglement, outperformance of standard learning-from-demonstration baselines, and more efficient and stable specialization.