acceptodds
Under review as a conference paper at ICLR 2027

When Should an AI Agent Change Its Mind? Counterfactual Evaluation of Multimodal Evidence Reasoning with Contrast-Persistence Tests

Abstract

Multimodal agents increasingly make decisions by acquiring evidence from heterogeneous sources, yet they are typically evaluated by final-answer accuracy alone. This leaves a basic reliability question unanswered: does an agent change its reported prediction when the evidence warrants reconsideration, while remaining stable when it does not? We introduce a counterfactual evaluation framework for multimodal evidence reasoning based on relevant addition, countervailing contradiction, exact redundancy, and important removal. The framework separates fixed-input response, autonomous whole-system response, matched-input conditional integration, and deterministic post-processing. This separation enables a contrast-persistence test: a whole-system difference should not be attributed to evidence integration when it changes materially after the compared integrators receive the same prior state and newly added evidence. Apparent robustness should likewise not be attributed to model reasoning when deterministic post-processing alone recreates it. ADNI remains our primary study, comparing Direct, Tool Agent, and Evidence Ledger pipelines across three open-weight backbones. Under a common-input, common-prior addition replay, the Ledger–Tool contrast, averaged equally across the tested backbones for each subject, changes from autonomously to 0.012 conditionally, a gap of (95% CI to ). In IBDome, the corresponding four-backbone contrast falls from 0.131 to 0.014, a gap of 0.117 (0.089–0.144). Exact-duplicate stability also degrades sharply when the Ledger preservation guard is removed and is recovered by guard-only processing. Together, these results show that endpoint performance, autonomous response, conditional integration, and engineered invariance are distinct evaluation objects. Multimodal agents should therefore be evaluated not only by what they predict, but also by whether an attributed evidence-revision property survives controls on the pipeline components that can produce it.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.