Decision-Sufficient Representations for Partially Observed Control
Abstract
A latent state for control must decide what to forget. Reconstruction spends capacity on every background detail; passive prediction discards any cue it cannot forecast, even one that changes the best action. A latent state should instead retain the distinctions that alter controlled reward consequences, and only those. We make this precise with the decision Hankel matrix, whose rows are histories and whose columns are reward-rooted, observation-contingent policy trees. Its rank is the minimal dimension of any exact linear representation of these consequences; a basis of its columns yields a decision code from which every trusted tree value, the optimal action values, and the complete optimal-action set are exact functions; and equality of its rows is the coarsest quotient that preserves every trusted return. At finite capacity, a taskwise decision rate–distortion objective admits a finite-sample oracle inequality and a deployment-regret guarantee, and its zero-regret rate equals the log-size of a minimum reachable action-profile cover, attained by an encoder that ignores nuisance entirely. In exact computations on finite POMDPs, the decision rank stays at while a decision-null nuisance inflates the belief span -fold, rises by exactly one for an action-changing cue that observation prediction aliases, and, among million encoders, regret first vanishes at four codes, reached by a single nuisance-free encoder. Compress what is decision-null, not what is hard to predict.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.