acceptodds
Under review as a conference paper at ICLR 2027

Staying Consistent, Staying on Target: Revisable References for Visual Reasoning

Abstract

Visual references enable vision-language models to reuse object-level information across successive questions. However, when an early localization is incorrect, reusing it can propagate the mistake through later reasoning, even when the resulting answers remain mutually consistent. Recovering from such errors requires revising the localization while preserving the object intended by the task. We address this requirement through revisable references, which separate task-defining information from the model's localization estimates. In a binary repair model, we characterize when rewards for historical agreement oppose task-beneficial corrections. This motivates Referent-Preserving Policy Optimization (), a reinforcement learning framework that learns reference revisions from their consequences for subsequent reasoning. jointly optimizes revision and execution: edits are assessed against retaining the original reference under the same executor, while continuations are compared within the same edited state. These comparisons connect both decisions to downstream task utility. Answer accuracy alone, however, does not establish whether the maintained reference has been repaired. We therefore design a controlled evaluation protocol that separately assesses reference recovery, preservation of the intended object, and answer correctness, using paired reference semantics and shared starting states under matched information, supervision, and computational budgets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.