CoRe: co-evolving reward models to mitigate latent reward hacking in video diffusion
This arXiv preprint identifies 'latent reward hacking'—where optimizing a generator against a fixed latent reward leads to high predicted scores but degraded perceptual and motion quality—caused by distributional escape. The authors introduce CoRe, a co-evolving framework that continually refits latent reward models on the generator's current samples while anchoring them to real-video preferences; experiments on Wan2.1-T2V-1.3B report consistent generation-quality improvements over the pretrained model and prior alignment methods and avoid the collapse seen with fixed-reward optimization.