Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

PlurPO has models simulate multiple stakeholder perspectives, cutting endorsement of harmful intent by 89% on average and reducing general-advice sycophancy from 17.8% to 8.0%

The work proposes Pluralistic Preference Optimization (PlurPO): given inputs describing interpersonal conflicts, the language model identifies and simulates the relevant stakeholders and is then trained to prefer and generate responses acceptable to all stakeholders, using only signals the model produces about its own outputs and no ground-truth labels; across four datasets and four model families, PlurPO substantially reduces social sycophancy compared with prior methods, for example reducing the endorsement rate by 89% on average on statements of intent to cause harm and closing the gap to human endorsement rates on general advice questions from 17.8% to 8.0%, and the preference dataset built for an 8B model transfers effectively to a 32B model.