acceptodds
Under review as a conference paper at ICLR 2027

AM I STILL COUPLED TO MY ENVIRONMENT? TASK-AGNOSTIC RUNTIME TELEMETRY FROM THE REINFORCEMENT-LEARNING INTERACTION LOOP

Abstract

After deployment, reward reports whether a reinforcement-learning agent is achieving its designed objective, but not necessarily whether the interaction producing that outcome remains characteristic of healthy operation. As a taskspecific scalar projection, reward can map different agent–environment interaction regimes to similar outcomes. We therefore ask whether the standard observation, action and next-observation triplet, (S, A, S′), can provide a complementary, reward-independent runtime reading. We introduce the interaction entropy budget C = H(S)+H(A)+H(S′) and its coupling share P = I(S,A; S′)/C as a firstperson, reward-independent reading of the realised MDP transition. This is an accounting view of standard Shannon quantities, not a new MDP or informationtheoretic formalism. Combined with the residual-uncertainty companions Hf , Hb and ΔH, it yields an individually calibrated runtime monitor that uses the current interaction stream under a fixed representation, without reward, a predefined fault model or policy internals. Across 21 frozen SAC and PPO agents in HalfCheetah, the monitor detected 69.0% of 168 perturbation trials, compared with 44.0% for matched windowed reward; the agent-clustered advantage was 25.0 percentage points (95% interval [14.9,35.1]), and each alarmed on 5 of 21 unperturbed controls. In a within-cohort ablation at the same five-control-alarm count, the budget normalisation detects 9 points more than raw mutual information and 18 more than the tested min-entropy normalisation. Across these perturbations, the monitor detects interaction-regime changes that the task-outcome projection can miss. It complements rather than replaces reward: it asks whether the agent is still interacting with its environment in the characteristic way established during healthy operation. Such telemetry could flag interaction drift or partial decoupling and help interpret whether a reward change accompanies a broader change in the closed loop. For deployed autonomous systems, the results show why tracking task outcomes alone can be insufficient for runtime monitoring: operators may also need to verify that the underlying control loop remains characteristic of healthy operation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.