acceptodds
Under review as a conference paper at ICLR 2027

VeriACT: Action-Grounded Verification for Efficient Interactive Visual Reasoning

Abstract

Vision-language models can reason effectively about static visual inputs, yet interactive visual reasoning remains challenging because errors in perception or action can propagate across subsequent decisions. An agent may fail to notice the effect of an action, repeat ineffective actions, or continue planning from stale beliefs, wasting costly environment interactions. We introduce VeriACT, a task-agnostic interaction framework that makes verification an explicit and consequential part of the agent loop. VeriACT requires each action to state a falsifiable expectation and, before the next action, verifies the resulting environment transition against that expectation. It combines model-based verification with controller-grounded visual signals, including frame-change and repetition detection, multi-frame transition analysis, click-target confirmation, and terminal-state verification. Verification outcomes can trigger replanning and conditional memory updates, allowing new evidence to directly modify the agent's subsequent behavior. We evaluate VeriACT on ARC-AGI-3 against a free-form control while holding the vision-language model, observation interface, memory, and evidence tools fixed. In a 13-game evaluation, VeriACT improves mean RHAE from to and reduces environment actions by % on levels successfully completed by both conditions, at the cost of increased inference-time computation. These results suggest that explicit, action-grounded verification can trade additional inference-time computation for greater task progress and fewer environment interactions, even when the underlying reasoning model remains unchanged.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.