PRISM-BENCH: Diagnosing Grounded and Faithful Multimodal Reasoning Beyond Final-Answer Accuracy
Abstract
Final-answer accuracy reports whether a vision-language model was right, but not whether its answer was grounded in the image, whether its stated reasoning reflects that evidence, or what the answer cost to produce. We present PRISM-BENCH, a benchmark of 1,000 expert-curated high quality questions spanning five under-represented reasoning categories, in which every item carries an idealized reasoning trace. Because the traces supply a per-item ground truth for the reasoning *process*, they enable measurements that an accuracy-only benchmark cannot support. Our central contribution is CORE, a cost-aware metric that charges each response against the length a human actually needed for *that* item rather than against a global token budget. CORE reorders 24.2% of all model pairs relative to accuracy while remaining uncorrelated with response length (), so it measures a construct neither accuracy nor mean length recovers. The traces further support two diagnostics: a perception probe that attributes roughly 85% of frontier-model failures to misperception, placing the bottleneck at grounding rather than inference; and trace alignment, which predicts answer correctness for all 25 systems tested (macro AUROC 0.66) and survives length control, so it captures reasoning content rather than verbosity. Across 25 open and closed systems including GPT-5 and Claude Opus 4.6, the best accuracy is 66.0% against an 83.9% human baseline, recent frontier gains fall on recognition rather than spatial grounding, and thinking variants generate 2.2–2.5 as many tokens as their instruct counterparts for lower accuracy. The optional CORE-PRM variant is critic-agnostic by design, and we validate it with two answer-blind frontier judges. The benchmark, traces, tags, and evaluation code will be released upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.