acceptodds
Under review as a conference paper at ICLR 2027

RecAgent: Reinforcing Referring Expression Comprehension through Agentic Tool-Augmentation

Abstract

Referring expression comprehension (REC) requires localizing the image region described by a free-form expression, and multimodal large language models (MLLMs) provide a flexible interface for this task while remaining prone to three recurring failures: they bind compositional constraints inconsistently, adopt a candidate before sufficient visual evidence is gathered, and translate a correct semantic decision into an imprecise box. We present RecAgent, an agentic framework that recasts REC as active evidence acquisition and optimizes it with tool-augmented reinforcement learning. RecAgent is trained in two stages. A supervised cold start on RefCoT-4K, a self-generated corpus of executable chain-of-thought trajectories, equips the policy to decompose an expression, hypothesize candidate referents, and verify each hypothesis. At inference, the policy queries a diagnostic tool library comprising an Expression Decomposer, a Visual Grounder, an Attribute Matcher, and a Spatial Reasoner, selecting which tools to call and in what order according to the ambiguity at hand. To make this behavior learnable, we propose a verification-gated composite reward that couples continuous IoU-weighted localization feedback with format validity, tool-usage quality, and a trajectory-efficiency term activated on verified-correct answers, rewarding trajectories that are short, correct, and well formed. Across standard (RefCOCO/+/g) and challenging generalization (gRefCOCO, Ref-L4) benchmarks, RecAgent attains state-of-the-art results, demonstrating the effectiveness and efficiency of active visual reasoning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.