acceptodds
Under review as a conference paper at ICLR 2027

Beyond Pixels: What Blind Replacement Measures in Visual Agents

Abstract

Visual agents gather evidence through crops and other intermediate observations. A common process score compares a judge's belief under a real image and a blind reference, then interprets the contrast as visual credit. We show that this interpretation mixes answer parameterization with the structure of a crop turn. On V*-Bench, annotated crops improve answer accuracy over matched distractors by 8.9–10.5 points, yet the tagged-sequence score places them below black replacements. Across 32 serialization, path, and judge settings, structural increments exceed pixel increments in 27 cases; their median magnitudes are 4.72 and 0.31 nats. We separate two choices hidden by the original score: the belief target and the credited part of the turn. Under a shared GRPO setup, option-posterior turn credit reaches 52.5% and 48.0% exact match on V*-Bench and HR-Bench-4K. Matching total reward mass retains nearly all of this performance, whereas uniform or shuffled placement does not. Pixel-level credit yields the strongest target retrieval, and its advantage remains after controlling for crop area. Blind replacement is therefore an operator-specific comparison, not a self-defining measure of visual credit.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.