acceptodds
Under review as a conference paper at ICLR 2027

Self-Anchoring and Behavioral Dissociation in Closed-Loop Multimodal Verification

Abstract

Unified multimodal models (UMMs) increasingly verify their own visual generations in closed-loop self-critique and refinement. We show that such verification is not attribution-neutral. When the same image is framed as self-generated rather than generated by another model, UMMs systematically shift verification toward their generation-time intent, a phenomenon we call attribution-conditioned self-anchoring. Across nine UMMs spanning six LLM backbone families, self-anchoring exhibits distinct behavioral regimes and persists across tasks and capability-controlled stratification. We localize this effect to compact, task-stable sets of attention-shift coordinates and establish their causal role through matched ablations. Removing the top-ranked set reduces anchoring by 43% in Bagel while nearly eliminating attribution discrimination under an AUROC-based reliability measure, far exceeding random controls. The same intervention attenuates both behavioral anchoring and confirmation amplification, linking them to a shared internal pathway. A dependency-structure analysis further predicts when verification produces confirmation amplification rather than corruption, explaining the reversal from prior text-RAG findings. Together, our results identify self-attribution as a hidden control variable in multimodal verification, challenging the assumption that a model can neutrally verify its own generations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.