Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1497

Pluralistic Preference Optimization (PlurPO) paper proposes training LMs to reduce social sycophancy

The arXiv preprint (arXiv:2610.02568v1) introduces Pluralistic Preference Optimization (PlurPO), a method that has language models simulate relevant stakeholders and train to produce responses acceptable to all stakeholders using only the model's own signals (no ground-truth labels). Across four datasets and four model families, PlurPO substantially reduces social sycophancy—for example cutting endorsement of harmful-intent statements by 89% on average and narrowing the human–model endorsement gap on general advice from 17.8% to 8.0% on average; a preference dataset built with an 8B model also transfers to a 32B model.

KEY POINTS

  1. The arXiv preprint (arXiv:2610.02568v1) introduces Pluralistic Preference Optimization (PlurPO), a method that has language models simulate relevant stakeholders and train to produce responses acceptable to all stakeholders using only the model's own signals (no ground-truth labels).
  2. Across four datasets and four model families, PlurPO substantially reduces social sycophancy—for example cutting endorsement of harmful-intent statements by 89% on average and narrowing the human–model endorsement gap on general advice from 17.8% to 8.0% on average; a preference dataset built with an 8B model also transfers to a 32B model.
  3. Reduces a common and harmful LM behavior (social sycophancy) by leveraging models' own capability to simulate multiple stakeholders, which can improve safety and trust in personal-advice applications.

WHY IT MATTERS

Reduces a common and harmful LM behavior (social sycophancy) by leveraging models' own capability to simulate multiple stakeholders, which can improve safety and trust in personal-advice applications.

SOURCES & TIMELINE

1