acceptodds
Under review as a conference paper at ICLR 2027

InterSelf: Learning World Models under Hidden Action Execution

Abstract

World models typically assume that the action selected by a policy is the action that causes the next transition. This assumption fails when an unobserved execution process repeats, drops, or remaps commands, causing the model to conflate execution dynamics with environmental dynamics. We formalize this mismatch as an action-fidelity problem and quantify the predictive information lost by conditioning only on intended actions. We introduce InterSelf, which augments DreamerV3 with a recurrent latent execution model combining categorical action remapping and temporal persistence. Its effective-action representation conditions both world-model learning and latent imagination, without executed-action labels or corruption identities. We show that this channel covers the evaluated execution failures, while its latent mixtures are identifiable only when different actions leave sufficiently separable effects. Across a 22-game Atari study with three training seeds, InterSelf achieves higher corrupted returns than DreamerV3 in 14 games; in seven action-sensitive games, it outperforms both DreamerV3 and the capacity-matched InterSelf-Identity control. Variation across the remaining games is consistent with the predicted boundary: explicit execution modeling need not help when action effects are weak or redundant. These results establish hidden action execution as a distinct source of world-model misspecification and provide a principled way to model it.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.