Learning the Observation Function in POMDPs: Identifiability through State Invariance
Abstract
Policies for partially observable Markov decision processes (POMDPs) act on state observations. The observation function governs the probabilities of emitting observations at each state. We study the following problem: Identify the unique probabilities of an unknown observation function from a dataset of observations generated under an unknown policy. However, the underlying observation function of a POMDP is not identifiable in general. We therefore define a class of observation functions with the key property that observation probabilities are identical across all states that can emit them, for example where the probabilities of sensor failures are invariant, that is, consistent across states. We call this class of observation functions state-invariant and show them to be identifiable. We provide an optimization-based approach to learn the observation function from data: we start with a natural formulation of the problem as a nonlinear program (NLP) and convexify the NLP toward a disciplined convex program (DCP) solvable in polynomial time. We prove that the solutions to the optimization problems are probably approximately correct (PAC) with respect to the dataset. Our experiments support the theoretical findings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.