Pluralistic Preference Optimization (PlurPO) paper proposes training LMs to reduce social sycophancy
The arXiv preprint (arXiv:2610.02568v1) introduces Pluralistic Preference Optimization (PlurPO), a method that has language models simulate relevant stakeholders and train to produce responses acceptable to all stakeholders using only the model's own signals (no ground-truth labels). Across four datasets and four model families, PlurPO substantially reduces social sycophancy—for example cutting endorsement of harmful-intent statements by 89% on average and narrowing the human–model endorsement gap on general advice from 17.8% to 8.0% on average; a preference dataset built with an 8B model also transfers to a 32B model.