acceptodds
Under review as a conference paper at ICLR 2027

Action-free Causal State Inference for Reinforcement Learning

Abstract

In partially observable environments, the state an agent learns from its history decides how well and how efficiently reinforcement learning succeeds, and causality gives that state its right target, the causal state, which groups histories by the future they lead to and helps the agent tell apart precisely the situations that matter, whereas predictive states may also keep irrelevant history. In this paper, we extend causal state learning to the action-free setting, as humans infer the plot of a film without acting in it, and prove that, under a mild condition on data collection, the causal state learned from observations and rewards alone loses nothing relevant to predicting the future and can be even more compact than its action-based counterpart. We further characterize how fast forgetting removes the rest, so its strength can be set automatically. Building on this, Action-free Causal State Inference (ACSI) injects noise into its recurrent memory, so that prediction preserves what the future requires while the rest decays, with the noise self-calibrated to a target forgetting rate rather than tuned per task. Empirically, ACSI recovers ground-truth causal states 9 AUC points better than action-conditioned, contrastive and bisimulation baselines, and is the only frozen representation to solve both control tasks that provably require memory, using 2.6–5.7X fewer frames than recurrent reinforcement learning. On real robot laser streams, it identifies revisited places 4 AUC points better than all learned baselines, offering a principled route to compact states from action-free data that make reinforcement learning more sample-efficient. The code of our method and environments will be released upon acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.