acceptodds
Under review as a conference paper at ICLR 2027

Target-Specific Distortion in State Supervision under Policy Swap

Abstract

History encoders for partially observable control are often trained on hidden-state labels logged by a privileged teacher that sees the state, then deployed with a policy that sees only past observations and actions. The labels are correct, but the teacher's actions also reveal what it saw, so training fits a posterior that this policy swap invalidates. Prior analyses treat the hidden state as a whole. We show that the swap distorts a prediction target exactly when, at histories that deployment reaches, the teacher's action likelihood averaged over hidden paths consistent with the history varies with the target's value, and that only the part of the teacher's action information about the target can distort it, although this part need not do so. The exact posterior fitted to the log scores better on teacher histories and worse at deployment, and an exact identity splits a trained predictor's gap into distortion and fitting error. In synthetic channels with fixed histories and action information, coupling the target to the teacher's actions alone reverses the learned risk ordering in every seed, and where this condition and per-target action information disagree, trained predictors follow the condition. On the RockSample benchmark, predictors trained on exact posterior labels show the deployment gap in every seed at every tested level of teacher privilege. At full privilege, fixed controllers acting on them earn more return with deployment-posterior labels, fitting error rather than exact distortion sets the gap's size, and validation on teacher histories favored early checkpoints of the distorted predictors, but only under a loss restricted to targets the history leaves undetermined. When the environment model is known or identifiable from the state-labeled log, relabeling with the deployment posterior removes the distortion. Known teacher propensities allow only a high-variance reweighted evaluation, and otherwise the remedy is history-only data collection.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.