acceptodds
Under review as a conference paper at ICLR 2027

Looking but Not Updating: Visual Belief Inertia in Embodied Multimodal Models

Abstract

An embodied assistant must judge a scene when its action history describes an outcome that the image contradicts. We study visual belief inertia: state answers that follow an incorrect action-conditioned history in independent single-image queries. VBI-Bench comprises 2,285 robot-state questions and two complementary controls: vary history while fixing the image, or vary the image while fixing all text. These controls reveal that strong performance under both correct and incorrect histories can coexist with poor discrimination between images. On Qwen3-VL-8B, revision-invariant fine-tuning (RI-FT) raises full-conflict accuracy from 14.4% to 98.3% on a 713-item history grid, yet answers both members of only 6.5% of 552 fixed-text image pairs correctly. RI-FT-CF adds image pairs with shared history, question, and options but different targets to the same answer-supervised objective. Across three seeds, both-correct pair accuracy reaches 77.5%, alongside 94.6% full-conflict and 98.3% correct-history accuracy. Component comparisons link the paired gain to the added examples and show how history variation changes the joint performance across conditions. The resulting diagnosis and training recipe make image changes an explicit part of evaluating history-resistant state judgments.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.