GAP-DPO: geometry-aligned pair selection improves personalized preference optimization
The arXiv preprint arXiv:2610.00061v1 introduces GAP-DPO, an iterative method that selects preference pairs based on gradient alignment between DPO updates and expected user-utility gradients. The paper analyzes how pair selection controls whether Direct Preference Optimization (DPO) acts as an error-corrective or reinforcement-like update under off-policy sampling, and reports that GAP-DPO improves stylistic fidelity, preference alignment, and generation quality on personalized text-generation benchmarks compared to standard DPO variants.