RESEARCH · RESEARCH · #1497
Pluralistic Preference Optimization (PlurPO) paper proposes training LMs to reduce social sycophancy
The arXiv preprint (arXiv:2610.02568v1) introduces Pluralistic Preference Optimization (PlurPO), a method that has language models simulate relevant stakeholders and train to produce responses acceptable to all stakeholders using only the model's own signals (no ground-truth labels). Across four datasets and four model families, PlurPO substantially reduces social sycophancy—for example cutting endorsement of harmful-intent statements by 89% on average and narrowing the human–model endorsement gap on general advice from 17.8% to 8.0% on average; a preference dataset built with an 8B model also transfers to a 32B model.
KEY POINTS
- The arXiv preprint (arXiv:2610.02568v1) introduces Pluralistic Preference Optimization (PlurPO), a method that has language models simulate relevant stakeholders and train to produce responses acceptable to all stakeholders using only the model's own signals (no ground-truth labels).
- Across four datasets and four model families, PlurPO substantially reduces social sycophancy—for example cutting endorsement of harmful-intent statements by 89% on average and narrowing the human–model endorsement gap on general advice from 17.8% to 8.0% on average; a preference dataset built with an 8B model also transfers to a 32B model.
- Reduces a common and harmful LM behavior (social sycophancy) by leveraging models' own capability to simulate multiple stakeholders, which can improve safety and trust in personal-advice applications.
WHY IT MATTERS
Reduces a common and harmful LM behavior (social sycophancy) by leveraging models' own capability to simulate multiple stakeholders, which can improve safety and trust in personal-advice applications.