Frozen Policy Mixtures: How Deployment Limits Feedback in Imitation Learning
Abstract
We study how deployment affects imitation learning from clean expert labels in finite policy classes containing the expert. For two stochastic policies with a uniform prior and Bayesian updates, frozen per-step mixing has worst-case cumulative squared trajectory Hellinger loss of iterated-logarithm order in episode length. Matching bounds identify a feedback mechanism: failure ends informative visits, limiting the next update. Sampling once per episode or updating after each label removes this horizon dependence. This law is specific to trajectory discrepancy; no cumulative return-gap bound uniform in rounds and stochastic classes holds even for one-step episodes. On our two-state construction, however, terminal-success rewards yield the same separation in cumulative return gaps: iterated-logarithmic for frozen mixing and constant for both alternatives. We also characterize batch updates and likelihood tempering. For deterministic classes, matching return-gap bounds resolve the proposed class-size improvement for Warm-Stagger; plurality voting achieves the proposed rates in expectation under its single-state query model. Controlled experiments validate the trajectory and return calculations, update choices, larger stochastic classes, and deterministic voting.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.