REFINE: Active Visual Reasoning with Spatially Indexed Evidence Refinement
Abstract
Active visual reasoning requires models to acquire fine-grained observations and retain them across successive decisions. Appending every crop grows the visual context even when an inspection revisits a previously observed region. We introduce REFINE, a reinforcement-learning framework that organizes acquired evidence into a persistent spatial state. Each region owns a fixed set of slots initialized from a coarse observation; higher-resolution inspection replaces those slots while preserving evidence elsewhere. A frozen answer-only probe measures the change in reference-answer support between initial and final states. This bounded, signed utility augments correctness and format rewards within Group Relative Policy Optimization, favoring useful refinement without rewarding inspection count. Evaluations span visual search, mathematical reasoning, chart understanding, and real-world perception. REFINE-7B achieves 85.26 on V* positional search and 69.2 on RealWorldQA, compared with 82.89 and 66.4 for the GRPO baseline. Its MME-RealWorld perception and reasoning scores are 68.60 and 50.37. Component ablations show complementary benefits from spatial storage and utility-aware learning. REFINE connects selective observation to persistent evidence through local replacement and outcome-grounded policy optimization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.