Discrepancy Is Not Advice Utility: On Privileged Multi-Agent Distillation
Abstract
When should decentralized agents imitate a privileged teacher? Reproducibility-based weighting favors teacher actions that a local auxiliary policy can predict, but predictability alone need not make an imitation update beneficial. In a two-agent cooperative game with an optimal teacher and an exact auxiliary, we show that discrepancy-weighted Jensen–Shannon divergence (JSD) imitation yields strictly lower team return than constant weighting after a sufficiently small update of equal joint-logit norm. The weighting shifts update mass toward the agent whose change is harmful to the current team. Yet the ordering reverses at the limiting policies under continuous-time pure JSD imitation. The counterexample thus rules out a general local-improvement guarantee for this weighting rule, without implying that it harms learning in general. Separately, exact-reference calibration shows how auxiliary capacity and fitting affect discrepancy at fixed information. A five-task neural study compares combined quality/discrepancy weighting with quality-only weighting and reference systems. Two comparisons match both selected imitation coefficients and learning rates; the others compare separately selected procedures. For the fixed teachers and selected procedures, combined weighting improves over RL on two warehouse tasks but shows no statistically resolved gain over quality-only weighting on any of the five tasks. These results distinguish two questions: whether an agent can reproduce a teacher's actions, and whether following those actions improves team return.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.