acceptodds
Under review as a conference paper at ICLR 2027

Counterfactuals Without Resets: Identifying POMDPs

Abstract

Many problems require estimating a world-dynamics model from interaction data where the observations do not uniquely reveal the underlying state. Our goal is to estimate time-homogenous discrete Partially Observable Markov Decision Processes (POMDPs) from a single trajectory of action and observations without any resets.We find that access to counterfactual samples provides sufficient conditions to learn a larger class of POMDPs.Counterfactual samples are samples from the joint distribution of all possible observations sequences obtained from executing multiple action sequences from the same state.In general, when gathering data from a single trajectory, it is not possible to gather counterfactual samples for a partially observable system.We therefore restrict our attention to the class of POMDPs that can be described as deterministic finite-state machines corrupted with bounded transition and observation noise. We show that deterministic machines contain anchor sequences, which can fully determine the underlying state when observed. We provide an algorithm, Counterfactual Successive Projection (CfactSP), that uses ideas from tensor decomposition and nonnegative matrix factorization to leverage these anchor sequences to identify the POMDP from a single reset-free data sequence. CfactSP is provably robust to a bounded number of corruptions in transition and observations. Empirical results suggest sample-complexity is substantially lower in practice than the guaranteed bound.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.