acceptodds
Under review as a conference paper at ICLR 2027

Think, Compress, Answer: Reinforcement Learning for Multimodal Grounding via Reasoning Compression

Abstract

Extended reasoning has substantially improved multimodal large language models (MLLMs) on tasks requiring composi- tional and relational inference. However, multimodal ground- ing imposes a distinct requirement: reasoning must ultimately converge on a target-specific compact representation that sup- ports precise localization. As reasoning trajectories grow longer, they often expand into broad scene narratives and peripheral relational details, which can dilute grounded per- ceptual evidence and lead to ambiguous or incorrect predic- tions. Motivated by the information bottleneck principle, we propose Think-Compress-Answer (TCA), a reinforcement learning framework that decomposes multimodal grounding into three phases: free-form reasoning, target-aware context compression, and answer generation conditioned on the com- pressed context. TCA is optimized through a multi-stage Group Relative Policy Optimization (GRPO) with progressive rewards that establish the structured output protocol, mini- mize reasoning-specific information while preserving target- relevant content, and directly improve localization accuracy. This staged design enables MLLMs to retain the benefits of extended reasoning while distilling it into the minimal con- text required for grounding. Experiments across six multi- modal grounding benchmarks show that TCA attains substan- tial zero-shot performance gains, demonstrating the effective- ness of reasoning compression

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.