acceptodds
Under review as a conference paper at ICLR 2027

PAVIR: When Should Multimodal Cascades Revise, Not Adopt?

Abstract

What cascades treat as progress, a stronger visual answer, rests on an untested assumption: that the expert's answer is the decision the task scores, even when a required claim can make Conflict correct. We test that assumption on public ScienceQA, MMStar, and MMMU-Pro items with aligned, conflicting, and missing claims, separating who sees the text from whether the expert writes a full answer. keeps the claim and the source requirements, calls an image-only expert only when the sources already disagree, and recomputes Answer, Conflict, or RequestEvidence with one rule. With the rule fixed, isolating the observer raises ScienceQA accuracy by 9.66 points, and the gain sits in the conflict cells; the same isolation lowers accuracy when every output must be an answer. On MMStar, disagreement-gated revision reaches 88.67% at 399/1,200 calls (61 repairs, 15 harms), from 84.83% with no calls; a 256-token 32B model that sees the claim reaches 87.25%, and 88.50% only after the same rule is applied (65 repairs, 21 harms; paired interval ), while calling all 494 eligible disagreements yields 88.58%. The comparison rule, not the ranking score or the size of the generator, is what makes the call useful, and only when Conflict is a legal output.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.