Agents stop reading each other's rewards: watching peers' behavior is enough to cooperate and split returns more evenly
Lead
Multi-agent reinforcement learning no longer needs peers' private rewards: each agent evaluates others' transitions with its own learned reward model, learns cooperation in Escape Room, Clean Up, and Commons Harvest, and often divides returns more evenly than agents that see the true rewards.
Story
An agent can assess another agent's situation from observed behavior and outcomes alone, without reading that agent's private reward signal, and still learn cooperation in three sequential social dilemmas. Earlier social-preference methods fed the other agents' rewards directly into the objective, and a reward is a private signal that rarely exists outside simulation. Across Escape Room, Clean Up, and Commons Harvest, the best self-referenced learner exceeds plain PPO in both collective return and Nash welfare, and matches the true-reward learners in Escape Room and Clean Up.
Each agent fits a model of its own reward on its own transitions, applies that model to a peer's transitions to estimate the peer's situation, and feeds the estimate into an unchanged social preference. The social preference itself is left as it was; only the peer-reward input it used to require is replaced. The estimate enters in two places: Self-Referenced Reward adds it to the learning reward, and Self-Referenced Advantage uses it to weight policy updates; in Clean Up, SRR-IA+V reaches a Nash welfare of 270 against 70 for the true-reward TR-EI.
Whether the estimate belongs in the reward or in the policy update depends on the social preference: a preference that compares agents suits the reward, while a purely benevolent one is safer in the policy update. Earlier work did not separate these two entry points, treating social preferences mainly as an intrinsic reward. SRR-IA+V is the better of the two methods in every game, while SRA-IA never escapes Escape Room; SRA-EI is better in Escape Room and Commons Harvest.
What to watch
A next step is to test agents whose rewards come from different events, where a model of one's own reward may no longer suffice, or to keep that model as a prior and correct it from what others do in an inverse-reinforcement-learning style. SRR needs no actor-critic, so self-referenced preferences could carry over to value-based learners. Letting policies see the record of past episodes, or communicate, could turn sharing into deliberate turn-taking. Judging another agent's situation from memory rather than a single window could make SRR robust to partial views.
In Commons Harvest the self-referenced learners fall behind the true-reward learners on collective return, and the reward model there is a small head on the policy's features, dominated by the large beam penalty. Nash welfare in that game needs care, since a self-scaling control that ignores the others reaches the highest welfare. Under partial observability, the reward placement with the look-ahead collapses while the policy-update placement keeps cooperating. In the heterogeneous-valuation experiment no learner divides by value, so whether self-reference would mislead an agent that otherwise acts on values remains open.
