Compressing What an Agent Can Steer: Control-Aware Structural Entropy
Abstract
A state abstraction for reinforcement learning should preserve what the agent can control, yet reward-free codelength objectives for compressing transition graphs score how predictable the dynamics are once actions are averaged out, not how far an agent can steer them. We prove that this gap is structural for structural entropy (SE), a random-walk codelength used to build hierarchical state abstractions: the SE codelength of a Markov decision process under any encoding tree is a function of the action-marginalised kernel alone, so no SE variant computed from the passive chain can distinguish two MDPs that differ only in their action-conditioned structure. We then define the minimal repair: the control gain at resolution , , the codelength saved when the block-level code may condition on actions. At the finest partition it equals stationary-averaged uniform-prior empowerment (exactly Klyubin–Salge empowerment in balanced deterministic MDPs), and it is monotone under coarsening. Each estimate is calibrated against a matched null, a within-state action shuffle that preserves the passive chain and all per-state action marginals exactly. On tabular MDP families where controllability is a construction knob independent of connectivity, the calibrated estimator recovers the constructed gain to 0.015 bits (Spearman , , in both families), while effective rank is blind where the passive chain is fixed (); the passive SE baseline breached its preregistered blindness threshold (, ), which a post-hoc 100-run extension attributes to sampling noise. In quotient planning on a checkerboard-inversion gridworld (gate failed) no reward-free partition objective we tested — passive, action-conditioned, or profile consistency — reliably yields plannable quotients, although plan-preserving partitions exist and the control-aware objective found one (retained 0.95 of optimal return, in 1 of 15 cells). On 15 learned agents both exploratory predictions hold in sign but not significantly. All gates and verdicts are reported verbatim.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.