When Reading Is Not Verifying: Executable Audits Decompose Verification Failures in Vision–Language Models
Abstract
Reliable verification of text in images requires more than reading the fields or stating the right rule: a model must also execute that rule correctly. We introduce ContraLedger, an executable record-verification benchmark with matched valid/invalid twins and stage-wise diagnostics of extraction, criterion fidelity, completion, arithmetic execution, and final judgment. Across six checkpoints, 191/200 selected cases that pass source/valid controls, exact transcription, and a separate rule-rejection probe still accept the invalid record. Stronger prompting and larger test-time budgets recover many of these failures, so this 191/200 rate is a conditional diagnostic rather than a ceiling on verification ability. On four development-held-out schemas, Qwen3-VL-8B and Qwen3.5-27B generate SMT-equivalent criteria on all 256 primary calls per checkpoint yet still make 15 and 26 completed decision errors. Trace replay localizes 39/41 displayed error traces to incorrect arithmetic execution and the remaining two to incorrect transcribed inputs; none is a case where the terminal label contradicts the model’s displayed check. An authored executable checker and a restricted model-generated program baseline substantially improve pair accuracy. Together, these results turn a broad “read-versus-reason” failure into an auditable decomposition of verification and motivate evaluating verification as a pipeline rather than a single end-task score.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.