Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

TOPIC · ENTITY #5494

multimodal large language models (MLLMs)

Related event timeline, sources and context from the news index.

EVENT TIMELINE

2

RESEARCH · 1 SOURCE · arXiv cs.AI

DEEPO: dual-entropy enhanced policy optimization to reduce hallucination in MLLMs

The paper (arXiv:2609.28570v1) proposes DEEPO, a dual-stage RL enhancement for multimodal LLMs that combines semantic-entropy-triggered expert prefixes to restore advantage variance and advantage-sign-aware Renyi preconditioning to counteract logit saturation. Both components individually outperform GRPO and together yield statistically significant gains on the VideoMMMU benchmark (+4.0, 95% CI [1.1, 6.9]); DEEPO reduces hallucination while preserving accuracy and training stability.

6.0

RESEARCH · 1 SOURCE · arXiv cs.AI

PolyBridgeBench: benchmarking MLLMs for physics-grounded bridge design

PolyBridgeBench is an executable benchmark that asks multimodal LLMs to produce complete node–member–material bridge topologies from a visual scene and engineering constraints, then runs deterministic legality checks and a native dynamic physics simulation. The benchmark reports deterministic validity, dynamic functional success, and post-failure recovery under a fixed interaction budget; experiments on six representative MLLMs across 189 levels reveal a large gap between deterministic validity and dynamic success, strong sensitivity to material budgets, and limited ability to repair after failures.

6.0