RESEARCH · RESEARCH · #1020
DEEPO: dual-entropy enhanced policy optimization to reduce hallucination in MLLMs
The paper (arXiv:2609.28570v1) proposes DEEPO, a dual-stage RL enhancement for multimodal LLMs that combines semantic-entropy-triggered expert prefixes to restore advantage variance and advantage-sign-aware Renyi preconditioning to counteract logit saturation. Both components individually outperform GRPO and together yield statistically significant gains on the VideoMMMU benchmark (+4.0, 95% CI [1.1, 6.9]); DEEPO reduces hallucination while preserving accuracy and training stability.
KEY POINTS
- The paper (arXiv:2609.28570v1) proposes DEEPO, a dual-stage RL enhancement for multimodal LLMs that combines semantic-entropy-triggered expert prefixes to restore advantage variance and advantage-sign-aware Renyi preconditioning to counteract logit saturation.
- Both components individually outperform GRPO and together yield statistically significant gains on the VideoMMMU benchmark (+4.0, 95% CI [1.1, 6.9]); DEEPO reduces hallucination while preserving accuracy and training stability.
- It targets two identified failure modes in RL-based fine-tuning—vanishing group-relative advantage on hard queries and gradient invisibility of confident-but-wrong tokens—offering a method to more effectively correct hallucinations in MLLMs.
WHY IT MATTERS
It targets two identified failure modes in RL-based fine-tuning—vanishing group-relative advantage on hard queries and gradient invisibility of confident-but-wrong tokens—offering a method to more effectively correct hallucinations in MLLMs.