DEEPO: dual-entropy enhanced policy optimization to reduce hallucination in MLLMs
The paper (arXiv:2609.28570v1) proposes DEEPO, a dual-stage RL enhancement for multimodal LLMs that combines semantic-entropy-triggered expert prefixes to restore advantage variance and advantage-sign-aware Renyi preconditioning to counteract logit saturation. Both components individually outperform GRPO and together yield statistically significant gains on the VideoMMMU benchmark (+4.0, 95% CI [1.1, 6.9]); DEEPO reduces hallucination while preserving accuracy and training stability.