What drives counterfactual-resoponse gains in policy distillation?
Abstract
Can teacher interpretability improve policy distillation beyond ordinary off-policy sampling? We study two routes: selecting a decision tree's input vocabulary and augmenting its training data with matched state pairs drawn along response-validated features. Across 27 PPO and SAC teachers on four control tasks, counterfactual augmentation improves intervention consistency (IC), a descriptive statistic of teacher–student response agreement, by on fresh pairs at equal tree size and a 20,000-label budget. Sampling controls recover most of this gain; the increment over random pairs is . A post-hoc expansion across 33 teachers evaluates IC and action error on identical endpoints: the scores are strongly associated, and only 58 of 2,376 method comparisons survive endpoint-error, return and competence matching, with a median absolute IC gap of . This does not establish equivalence, but provides no demonstrated incremental diagnostic utility for IC. Restricting the vocabulary to validated features can omit information needed for control; a predictive sufficiency check repairs the Pendulum failure without establishing a general repair. The study combines a pre-registered core with explicitly post-hoc controls and sensitivities. Its contribution is an empirical decomposition of response gains and the limits of this gate and feature library, without an established interpretability-specific control or robustness advantage.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.