RESEARCH · RESEARCH · #531
Learning Heterogeneous Preferences (arXiv:2609.17847v1)
This preprint introduces "individuated utility" models that condition utility on both the individual and decision context and proposes a multi-stage architecture to estimate them from multi-modal data. The authors evaluate on a new dataset of 575,000 pairwise aesthetic judgments from 2,398 participants comparing automotive wheel designs and report that individuated utility models substantially outperform universal utility and foundation-model baselines, arguing that annotator disagreement reflects systematic preference heterogeneity rather than noise.
KEY POINTS
- This preprint introduces "individuated utility" models that condition utility on both the individual and decision context and proposes a multi-stage architecture to estimate them from multi-modal data.
- The authors evaluate on a new dataset of 575,000 pairwise aesthetic judgments from 2,398 participants comparing automotive wheel designs and report that individuated utility models substantially outperform universal utility and foundation-model baselines, arguing that annotator disagreement reflects systematic preference heterogeneity rather than noise.
- Accounting for systematic individual preference heterogeneity can improve reward models and human-aligned policy learning by making clear whose preferences a model represents.
WHY IT MATTERS
Accounting for systematic individual preference heterogeneity can improve reward models and human-aligned policy learning by making clear whose preferences a model represents.