Comedic Fool's Gold: reward exploits and countermeasures in conversational humor (arXiv:2610.00197v1)
This new arXiv preprint studies automated reward functions for training language models to produce conversational humor, finding multiple exploit paths and evaluating fixes. An embedding-based surprise reward accepts word-shuffled replies as well as witty ones (partly caught by a fluency filter that nonetheless sometimes rejects valid jokes), an audience-model predictor is vulnerable to laughter cues which can be mitigated by normalizing cues across speakers but still leaves other exploits, and three reinforcement-learning training runs with successive reward revisions improved a combined evaluation score by 0.0903 and reduced zero-score sessions by 40% while failing to meet a preregistered humor-specific improvement target.