Identifiability and exploration in performative reinforcement learning
Abstract
In performative reinforcement learning, deploying a policy changes the environment it acts in. The usual algorithms retrain on the environment the current policy induces and repeat, converging to a performatively stable policy. They do not explore, so their data come only from the state–action pairs their own policies keep visiting, and we ask what those data can identify. For performative MDPs whose response is an exponential tilt of a base model, the answer is exact in the limit: the data identify the response directions revealed by the pairs that stay visited, and nothing else. The directions left over form a blind subspace that can be computed before any data arrive, from a candidate policy and the known response design. This separates two cases. When no blind direction can change which policy is best, retraining loses coverage but not identifiability, as on the grid-world benchmark used in this literature. When one can, retraining can settle on a policy that its own data cannot refute, and a learner whose cumulative exposure to the revealing pairs stays bounded has linear regret. We analyse three remedies with regret guarantees: optimism and posterior sampling, which explore only among plausibly optimal policies yet suffice once rewards are observed, and minimal probing, which occasionally plays a cheap set of revealing pairs and needs only the planner that retraining already uses. Where the test fails, probing cuts the regret of certainty equivalence by 91% on generated environments and by 61% in a simulated battery charger whose lithium plating is hidden below an onset, with no detected difference from optimism at a quarter of its computation. In environments drawn without a blind spot built in, the failure is common above a dimension threshold and absent below it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.