Target-Specific Distortion in State Supervision under Policy Swap
Abstract
History encoders for partially observable control are often trained on hidden-state labels logged by a privileged teacher that sees the state, then deployed with a policy that sees only past observations and actions. The labels are correct, but the teacher's actions also reveal what it saw, so training fits a posterior that this policy swap invalidates. Prior analyses treat the hidden state as a whole. We show that the swap distorts a prediction target exactly when, at histories that deployment reaches, the teacher's action likelihood averaged over hidden paths consistent with the history varies with the target's value, and that only the part of the teacher's action information about the target can distort it, although this part need not do so. The exact posterior fitted to the log scores better on teacher histories and worse at deployment, and an exact identity splits a trained predictor's gap into distortion and fitting error. In synthetic channels with fixed histories and action information, coupling the target to the teacher's actions alone reverses the learned risk ordering in every seed, and where this condition and per-target action information disagree, trained predictors follow the condition. On the RockSample benchmark, predictors trained on exact posterior labels show the deployment gap in every seed at every tested level of teacher privilege. At full privilege, fixed controllers acting on them earn more return with deployment-posterior labels, fitting error rather than exact distortion sets the gap's size, and validation on teacher histories favored early checkpoints of the distorted predictors, but only under a loss restricted to targets the history leaves undetermined. When the environment model is known or identifiable from the state-labeled log, relabeling with the deployment posterior removes the distortion. Known teacher propensities allow only a high-variance reweighted evaluation, and otherwise the remedy is history-only data collection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.