Towards a Theoretical Understanding of How Mutual Distillation Improves Generalization
Abstract
We argue that mutual distillation between reinforcement learning policies serves as implicit regularization that prevents them from overfitting to irrelevant features. Theoretically, we establish the first trust region lower bound connecting generalization performance with policy robustness, confirming the long-standing assumption that policy robustness underpins generalization. We further characterize the missing connection between deep mutual learning (DML) and this lower bound. Empirically, we demonstrate that mutual distillation between policies promotes such robustness, enabling the spontaneous emergence of invariant representations from visual inputs. Ultimately, rather than emphasizing extensive baseline comparisons, our goal is to uncover the underlying principles of generalization and provide theoretical insights into its mechanisms as our foundational contributions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.