Surprise Is Not Enough: Dynamics-Consistent Observation Reliability in Reinforcement Learning
Abstract
Reinforcement learning agents deployed on physical systems act on what their sensors report, and some sensor failures produce readings that still look plausible: a channel freezes at its last value, lags behind the true state, or drifts slowly away from it. A policy cannot distinguish such a reading from a real change in state, and robustness training does not tell it which observation to distrust. We introduce dynamics-consistent reliability (DCR), a per-step, per-channel estimate of whether an observation can be trusted, computed inside the RL loop from the agent's own experience. DCR compares each reading with what a learned dynamics model predicts from past observations and actions, and scores the residual two-sided on the model's own clean scale over several window lengths, the classical innovation test of fault diagnosis, so that a residual that is too small or too biased is itself evidence of a fault. The scoring rule has no fitted parameters and each of its components is standardized against a known null distribution. This design addresses a failure shared by the detectors used in practice, from temporal-jump thresholds and reconstruction error to ensemble disagreement and innovation gating: they score an observation by how surprising it is, and a plausible fault is not surprising; the temporal-jump detector ranks a frozen sensor as more reliable than a working one, with per-channel AUROC of 0.004–0.021 on three of four MuJoCo tasks, and the other surprise scores are weakened. We evaluate DCR on nine sensor-fault families with per-step ground truth, in closed loop, with every family held out of tuning when tested. Over five seeds DCR reaches mean held-out AUROC of 0.994 on Hopper, 0.986 on Walker2d, 0.995 on HalfCheetah and 0.792 on Ant, and outperforms PEDM- and DEXTER-style RL out-of-distribution detectors in every task; averaging the PEDM-style score over a window recovers most of that margin, so temporal accumulation is the larger lever and two-sided scoring adds a consistent 0.1–3.6 points. A learned head trained on injected faults is below the fixed statistic on average over held-out families. We further characterize when history-only estimation is impossible, derive a single-channel drift-detectability bound and show that the other channels carry most of the evidence, and build a controller that dead-reckons a flagged channel through the dynamics model and recovers 47–65% of the return lost to a fault on HalfCheetah, while failing on Walker2d, where isolation errs, and on two-channel faults. Code, fault definitions, trained policies and results are released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.