acceptodds
Under review as a conference paper at ICLR 2027

DEEPO: Dual-Entropy Enhanced Policy Optimization for Hallucination in MLLMs

Abstract

Reinforcement learning (RL) is widely used to sharpen reasoning in multimodal large language models (MLLMs), yet its effect on hallucination is uneven. We trace this to two weak points in the correction chain from reward to parameter update. At the rollout level, hard queries—those with high semantic entropy— frequently produce unanimously wrong sample groups, collapsing the group- relative advantage to zero exactly where hallucination risk is highest. At the optimization level, confident-but-wrong tokens are gradient-invisible: a categor- ical policy’s expected score-gradient norm vanishes as its distribution sharpens, so the predictions that most need correction receive the weakest updates. We propose Dual-Entropy Enhanced Policy Optimization (DEEPO), a dual-stage en- hancement combining signal variance regularization with gradient precondition- ing: semantic-entropy-triggered expert prefixes inject grounded continuations on high-uncertainty queries, providing direct supervision and restoring advantage variance, while advantage-sign-aware R´enyi preconditioning counteracts logit- level saturation so correction reaches confident errors in the operational confi- dence regime. Both branches improve over GRPO individually; their interaction is statistically significant on VideoMMMU—the most complex long-horizon task in our evaluation suite (+4.0, 95% CI [1.1, 6.9])—and additive elsewhere. DEEPO reduces hallucination while preserving accuracy and training stability.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.