acceptodds
Under review as a conference paper at ICLR 2027

Measuring the Local Action Utility of Visual History in Dreamer

Abstract

An agent can see the same image yet choose a different action because of its visual history. Does that change help it succeed? We measure this link by holding the current image fixed, changing past images, and evaluating the resulting first-action probabilities with the same continuing agent. A second measurement asks whether different current cues lead to different later outcomes when the first action is held fixed. Both use real simulator continuations and an explicit pairing of the recurrent policy's stochastic states. Controlled tasks distinguish information that is stored, information that changes actions, and information that improves outcomes. In Craftax-Classic, an exploratory DreamerV3 cohort motivates a prespecified confirmation on five new training seeds. Recent history improves local target success through the first-action choice by 5.13 and 5.36 percentage points under two masked continuation conditions. Every new model has a positive estimate, and joint model/world intervals remain positive after adjustment for three primary endpoints. An exploratory analysis links this gain to a shift toward the nominated tree; any-wood and native-reward gains remain unresolved. The separate continuation-square endpoint is also unresolved. The confirmed benefit concerns recent history with a visible current target, erased-history recipients, and masked later observations. A full-panel extension with clean recipient history and unmasked future images leaves its benefit unresolved. The measurements connect visual input to local action utility under explicit execution conditions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.