acceptodds
Under review as a conference paper at ICLR 2027

RewardLens: Measuring the Behavioral Resolution of Static Multimodal Judge Scores

Abstract

Static preference accuracy is a common way to compare multimodal judges. A tie shows that two judges were equally accurate on the evaluated examples, but says little about how they respond when decision-relevant visual evidence changes. We call the extent to which a static score constrains such intervention behavior its behavioral resolution. RewardLens measures this property with matched visual triplets that hold the question and candidate responses fixed: relevant edits reverse the programmed gold preference, while matched irrelevant edits preserve it. We quantify the two responses with Relevant Adaptation (RA) and Irrelevant Invari- ance (II). On an independent same-domain CLEVR static evaluation, the same five judges score 100% on both Attribute and Spatial, yet their RA spans are 0 and 43.5 percentage points. In Count, Qwen and Skywork score 100.0% and 99.5% statically but differ by 45.73 points in RA on shared correct bases while remaining nearly perfectly invariant to irrelevant edits. Across all 42 frozen one-point same- domain matches, 16 have zero RA gap and 16 have gaps of at least 10 points. Static accuracy summarizes correctness on fixed examples. When responses to changing evidence matter, that summary should be paired with controlled inter- vention measurements.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.