RESEARCH · RESEARCH · #942
Predicting objective conflict and covering trade-offs for steerable pluralistic alignment (arXiv:2609.26929v1)
This paper studies steerable pluralistic alignment using Multi-Objective Direct Preference Optimization (MODPO). Across seven objective pairs drawn from HelpSteer and UltraFeedback, the authors ask when a single model can improve two objectives simultaneously and how to cover many trade-offs without training separate models; they find two pre-training measurements predict whether objectives align or conflict for human-annotated data (but not for AI-annotated data, where response length and repetition confound reward-model scores), and that selecting the nearest trained model or merging parameters helps expand trade-off coverage though neither reliably matches direct training.
KEY POINTS
- This paper studies steerable pluralistic alignment using Multi-Objective Direct Preference Optimization (MODPO).
- Across seven objective pairs drawn from HelpSteer and UltraFeedback, the authors ask when a single model can improve two objectives simultaneously and how to cover many trade-offs without training separate models; they find two pre-training measurements predict whether objectives align or conflict for human-annotated data (but not for AI-annotated data, where response length and repetition confound reward-model scores), and that selecting the nearest trained model or merging parameters helps expand trade-off coverage though neither reliably matches direct training.
- The paper gives practical diagnostics and methods for building steerable models that balance conflicting human preferences and shows limitations of reward-model signals and transfer methods when covering many trade-offs.
WHY IT MATTERS
The paper gives practical diagnostics and methods for building steerable models that balance conflicting human preferences and shows limitations of reward-model signals and transfer methods when covering many trade-offs.