NEWS · RESEARCH · #134
Safety as a Constraint: Fine-tuning an LLM Recommender to Explain Itself
This arXiv preprint describes fine-tuning a recommender LLM to generate personalized, faithful, and strictly non-harmful explanations for video recommendations using two LLM-based judge reward models and a constrained GRPO training procedure. On a held-out real-world test set the tuned model raises the all-three-criteria PASS rate from 0.649 to 0.956 under the authors' judges and from 0.677 to 0.931 under an independent judge, while the frontier generator performs similarly to the untuned recommender and the model's recommendation and language abilities remain unchanged.
KEY POINTS
- This arXiv preprint describes fine-tuning a recommender LLM to generate personalized, faithful, and strictly non-harmful explanations for video recommendations using two LLM-based judge reward models and a constrained GRPO training procedure.
- On a held-out real-world test set the tuned model raises the all-three-criteria PASS rate from 0.649 to 0.956 under the authors' judges and from 0.677 to 0.931 under an independent judge, while the frontier generator performs similarly to the untuned recommender and the model's recommendation and language abilities remain unchanged.
- It suggests LLM-based recommenders can be trained to produce faithful, non-harmful explanations without harming recommendation quality, informing safer, unified agentic interfaces.
WHY IT MATTERS
It suggests LLM-based recommenders can be trained to produce faithful, non-harmful explanations without harming recommendation quality, informing safer, unified agentic interfaces.