acceptodds
Under review as a conference paper at ICLR 2027

Learning the Observation Function in POMDPs: Identifiability through State Invariance

Abstract

Policies for partially observable Markov decision processes (POMDPs) act on state observations. The observation function governs the probabilities of emitting observations at each state. We study the following problem: Identify the unique probabilities of an unknown observation function from a dataset of observations generated under an unknown policy. However, the underlying observation function of a POMDP is not identifiable in general. We therefore define a class of observation functions with the key property that observation probabilities are identical across all states that can emit them, for example where the probabilities of sensor failures are invariant, that is, consistent across states. We call this class of observation functions state-invariant and show them to be identifiable. We provide an optimization-based approach to learn the observation function from data: we start with a natural formulation of the problem as a nonlinear program (NLP) and convexify the NLP toward a disciplined convex program (DCP) solvable in polynomial time. We prove that the solutions to the optimization problems are probably approximately correct (PAC) with respect to the dataset. Our experiments support the theoretical findings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.