Not All Tokens Contribute Equally: Token-Level Reinforcement Learning for Multimodal Hallucination Mitigation
Abstract
Reinforcement learning with verifiable rewards (RLVR) is increasingly used to reduce hallucination in multimodal large language models (MLLMs). A key limitation is that the RLVR reward is defined only at the sequence level, so every token in a response shares the same advantage, without explicitly distinguishing its contribution to the error. Two common failure modes motivate our approach. The model may attend insufficiently to the image and rely on language priors, or make a wrong choice at an uncertain reasoning step. Neither failure mode is directly observable, but the model provides two relevant internal signals during training. The visual attention ratio of a token provides a proxy for visual grounding, and the predictive entropy indicates where the model is uncertain. Building on these two signals, we take a token-level view of RLVR and propose **VIEW** (**V**isual-grounding and **I**nformation-**E**ntropy **W**eighting), which reshapes the learning signal at the token level. VIEW first rescales each token's advantage using visual attention, then focuses the policy gradient on high-entropy tokens where reasoning may fork. Both signals are used only during training, and VIEW leaves standard autoregressive inference unchanged. Extensive experiments across hallucination, visual-reasoning, and general multimodal benchmarks show that VIEW improves the Qwen3-VL-8B average over GRPO and PAPO.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.