Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1020

DEEPO: dual-entropy enhanced policy optimization to reduce hallucination in MLLMs

The paper (arXiv:2609.28570v1) proposes DEEPO, a dual-stage RL enhancement for multimodal LLMs that combines semantic-entropy-triggered expert prefixes to restore advantage variance and advantage-sign-aware Renyi preconditioning to counteract logit saturation. Both components individually outperform GRPO and together yield statistically significant gains on the VideoMMMU benchmark (+4.0, 95% CI [1.1, 6.9]); DEEPO reduces hallucination while preserving accuracy and training stability.

KEY POINTS

  1. The paper (arXiv:2609.28570v1) proposes DEEPO, a dual-stage RL enhancement for multimodal LLMs that combines semantic-entropy-triggered expert prefixes to restore advantage variance and advantage-sign-aware Renyi preconditioning to counteract logit saturation.
  2. Both components individually outperform GRPO and together yield statistically significant gains on the VideoMMMU benchmark (+4.0, 95% CI [1.1, 6.9]); DEEPO reduces hallucination while preserving accuracy and training stability.
  3. It targets two identified failure modes in RL-based fine-tuning—vanishing group-relative advantage on hard queries and gradient invisibility of confident-but-wrong tokens—offering a method to more effectively correct hallucinations in MLLMs.

WHY IT MATTERS

It targets two identified failure modes in RL-based fine-tuning—vanishing group-relative advantage on hard queries and gradient invisibility of confident-but-wrong tokens—offering a method to more effectively correct hallucinations in MLLMs.

SOURCES & TIMELINE

1