CXR-ContraBench: Auditing Polarity Reversals Beyond Aggregate Accuracy
Abstract
A medical vision-language model can identify a finding on a chest X-ray and then deny it after a single answer option changes. We call this directional failure *negated-option attraction*. CXR-ContraBench measures it with a paired protocol that holds the image, question, reference finding, and retained options fixed while one distractor is replaced by the negated target, and records every answer transition, including correct answers that become false denials and wrong answers that recover. On 260 paired CheXpert items, MedGemma-4B-it turns 89 of 158 initially correct answers into false denials and Qwen2.5-VL-7B turns 41 of 116, conditional reversal rates of 56.33% and 35.34%. The aggregate accuracy of Qwen2.5-VL nevertheless rises by 5.77 percentage points because recoveries offset these reversals. Because the target-paired answer set admits an image-free structural ceiling, image-conditioned validity is evaluated separately: matched, opposite-label, shuffled, and blank images change which statement is true while the text stays fixed. Under this diagnostic, expected polarity updating is limited, a polarity-explicit prompt changes answer behavior with a benefit that depends on the measured endpoint, and in-format adaptation reaches perfect format accuracy without paired image-conditioned correctness. The transition persists across ten wording families, option-free single-sentence generation, OpenI, frontier models, and 135,754 CheXpert records from 59,173 studies. Average accuracy should therefore be interpreted together with whether previously correct judgments are preserved. Code is available at https://anonymous.4open.science/r/cxr-contrabench/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.