acceptodds
Under review as a conference paper at ICLR 2027

BlameBench: Diagnosing Atomic Perception in Multimodal Reasoning Benchmarks

Abstract

Multimodal large language models ground reasoning in fine-grained visual evidence, but final-answer accuracy conflates errors in perception and subsequent reasoning. We introduce **BlameBench**, a task-grounded benchmark that uses visual evidence from solution trajectories with correct final answers to decide what to test. We extract visual observations, verify them against source images, and convert the resulting facts into atomic perception questions that can be answered without solving the full source problem. This construction yields a fixed 2,500-question benchmark spanning six source benchmarks and seven reporting categories, including 2,058 direct-perception questions and 442 supplementary concept probes. Across 11 frontier models, source and category profiles reveal differences concealed by overall accuracy: Qwen 3.8 Max and Claude Opus 5 are separated by only 0.64 percentage points overall but lead on different sources and categories. The same questions also provide an upstream readout: across four InternVL3.5-Pretrained checkpoints, source-centered scores before fine-tuning correlate with downstream accuracy after fine-tuning (). **BlameBench** thus supports final-model perceptual diagnosis and pretraining-stage diagnosis.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.