Beyond Pixels: What Blind Replacement Measures in Visual Agents
Abstract
Visual agents gather evidence through crops and other intermediate observations. A common process score compares a judge's belief under a real image and a blind reference, then interprets the contrast as visual credit. We show that this interpretation mixes answer parameterization with the structure of a crop turn. On V*-Bench, annotated crops improve answer accuracy over matched distractors by 8.9–10.5 points, yet the tagged-sequence score places them below black replacements. Across 32 serialization, path, and judge settings, structural increments exceed pixel increments in 27 cases; their median magnitudes are 4.72 and 0.31 nats. We separate two choices hidden by the original score: the belief target and the credited part of the turn. Under a shared GRPO setup, option-posterior turn credit reaches 52.5% and 48.0% exact match on V*-Bench and HR-Bench-4K. Matching total reward mass retains nearly all of this performance, whereas uniform or shuffled placement does not. Pixel-level credit yields the strongest target retrieval, and its advantage remains after controlling for crop area. Blind replacement is therefore an operator-specific comparison, not a self-defining measure of visual credit.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.