Public articles linked to the same research event.
arXiv The work proposes Pluralistic Preference Optimization (PlurPO): given inputs describing interpersonal conflicts, the language model identifies and simulates the relevant stakeholders and is then trained to prefer and generate responses acceptable to all stakeholders, using only signals the model produces about its own outputs and no ground-truth labels; across four datasets and four model families, PlurPO substantially reduces social sycophancy compared with prior methods, for example reducing the endorsement rate by 89% on average on statements of intent to cause harm and closing the gap to human endorsement rates on general advice questions from 17.8% to 8.0%, and the preference dataset built for an 8B model transfers effectively to a 32B model.
The work proposes Pluralistic Preference Optimization (PlurPO): given inputs describing interpersonal conflicts, the language model identifies and simulates the relevant stakeholders and is then trained to prefer and generate responses acceptable to all stakeholders, using only signals the model produces about its own outputs and no ground-truth labels; across four datasets and four model families, PlurPO substantially reduces social sycophancy compared with prior methods, for example reducing the endorsement rate by 89% on average on statements of intent to cause harm and closing the gap to human endorsement rates on general advice questions from 17.8% to 8.0%, and the preference dataset built for an 8B model transfers effectively to a 32B model.
The work proposes Pluralistic Preference Optimization (PlurPO): given inputs describing interpersonal conflicts, the language model identifies and simulates the relevant stakeholders and is then trained to prefer and generate responses acceptable to all stakeholders, using only signals the model produces about its own outputs and no ground-truth labels; across four datasets and four model families, PlurPO substantially reduces social sycophancy compared with prior methods, for example reducing the endorsement rate by 89% on average on statements of intent to cause harm and closing the gap to human endorsement rates on general advice questions from 17.8% to 8.0%, and the preference dataset built for an 8B model transfers effectively to a 32B model.
The work proposes Pluralistic Preference Optimization (PlurPO): given inputs describing interpersonal conflicts, the language model identifies and simulates the relevant stakeholders and is then trained to prefer and generate responses acceptable to all stakeholders, using only signals the model produces about its own outputs and no ground-truth labels; across four datasets and four model families, PlurPO substantially reduces social sycophancy compared with prior methods, for example reducing the endorsement rate by 89% on average on statements of intent to cause harm and closing the gap to human endorsement rates on general advice questions from 17.8% to 8.0%, and the preference dataset built for an 8B model transfers effectively to a 32B model.