acceptodds
Under review as a conference paper at ICLR 2027

Deeper Is Not Always Wiser: Evidential GRPO for Overthinking in Large Vision-Language Models

Abstract

Existing approaches to overthinking in large vision-language models (LVLMs) mainly control reasoning length at the token-sequence level, with limited attention to how entropy changes across LLM depth. We observe that entropy can reach a trough at intermediate layers but rebound as LLM depth increases, leading to unnecessary additional reasoning. Motivated by this observation, we propose Evidential GRPO EviGRPO, a new GRPO variant that regulates reasoning through a Trigger–Verify process. It consists of an entropy trigger, which monitors entropy variations across LLM layers to identify the layer where entropy rebounds from its trough, and an evidence verifier, which detects rebound-induced token preference shifts and learns to assess whether the shift is sufficiently supported by evidence. For an evidence-unsupported preference shift, the final-layer token is replaced with the token predicted at the entropy-trough layer, while evidence-supported tokens are retained. Next, we formulate the EviGRPO reward based on the changes in reasoning length and final-answer correctness under trough-token substitution, while deriving evidential supervision from token-shift outcomes for verifier optimization. In this way, EviGRPO rectifies abnormal final-layer token shifts, thereby mitigating their propagation into redundant reasoning. Experiments across seven reasoning benchmarks demonstrate that EviGRPO consistently improves reasoning accuracy across 3B, 7B, and 32B models, achieving accuracy gains of 1.2–5.7 percentage points over standard GRPO while reducing reasoning length by up to 256.4 tokens, effectively balancing reasoning accuracy and efficiency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.