RESEARCH · RESEARCH · #1233
CoRe: co-evolving reward models to mitigate latent reward hacking in video diffusion
This arXiv preprint identifies 'latent reward hacking'—where optimizing a generator against a fixed latent reward leads to high predicted scores but degraded perceptual and motion quality—caused by distributional escape. The authors introduce CoRe, a co-evolving framework that continually refits latent reward models on the generator's current samples while anchoring them to real-video preferences; experiments on Wan2.1-T2V-1.3B report consistent generation-quality improvements over the pretrained model and prior alignment methods and avoid the collapse seen with fixed-reward optimization.
KEY POINTS
- This arXiv preprint identifies 'latent reward hacking'—where optimizing a generator against a fixed latent reward leads to high predicted scores but degraded perceptual and motion quality—caused by distributional escape.
- The authors introduce CoRe, a co-evolving framework that continually refits latent reward models on the generator's current samples while anchoring them to real-video preferences; experiments on Wan2.1-T2V-1.3B report consistent generation-quality improvements over the pretrained model and prior alignment methods and avoid the collapse seen with fixed-reward optimization.
- CoRe matters because making reward models co-evolve with the generator addresses distributional escape and reduces latent reward hacking, improving alignment and generation quality in video diffusion systems.
WHY IT MATTERS
CoRe matters because making reward models co-evolve with the generator addresses distributional escape and reduces latent reward hacking, improving alignment and generation quality in video diffusion systems.