acceptodds
Under review as a conference paper at ICLR 2027

Less Is More: Selective Expert Supervision for Multi-Teacher On-Policy Distillation

Abstract

Domain-specialist language-model agents can perform well individually, but integrating their capabilities into one student remains difficult. Multi-Teacher On-Policy Distillation (MOPD) addresses sequential-training interference by supervising student rollouts with domain specialists, yet our four-domain study reveals uneven transfer: WebShop success is 57.29% for its teacher and only 40.53% for the Vanilla MOPD student. A directed gradient analysis of Vanilla MOPD further finds time-varying opposition among domain learning signals, raising the question of whether every student trajectory needs full teacher supervision. We propose SEAM (Selective Expert Supervision and Advantage Matching), which preserves rollout quotas and fixed teacher routing but scores every failed trajectory and only a sampled subset of successful ones. To limit changes in domain influence when retained batches shrink, SEAM calibrates each domain's distillation signal to its original rollout budget before a shared student update. Across ALFWorld, WebShop, SciWorld, and Search, SEAM improves WebShop success by 9.35 percentage points and four-domain macro success by 2.20 points over Vanilla MOPD across two training seeds. It uses 5.75% fewer teacher-scored trajectories and reduces measured end-to-end training time by 5.48%. Selective supervision thus improves integration without reducing exploration, although specialist performance is not inherited losslessly.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.