MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate
Abstract
On-policy distillation (OPD) trains a student on its own trajectories with token-level teacher supervision, but existing methods face a single-teacher capability ceiling: when the teacher and student make the same mistake, the teacher can reinforce it. This problem is sharper in agentic tasks, where per-step errors compound across long trajectories. We introduce MAD-OPD (Multi-Agent Debate-driven On-Policy Distillation), which uses debate among teachers to improve supervision at states visited by the student. The teachers use the debate transcript to evaluate the student's tokens and report confidence scores that determine their loss weights. For agentic tool use, we introduce On-Policy Agentic Distillation (OPAD), which performs step-wise OPD by sampling each tool action from the student and applying token-level supervision at that decision state. Our analysis of gradients and teacher aggregation motivates Jensen–Shannon divergence for agentic tasks and reverse Kullback–Leibler divergence for code generation. The divergence ablation supports these choices in the tested configuration. Across six teacher–student configurations spanning Qwen3 and Qwen3.5, 1.7B–14B students, and 8B–32B teachers on five agentic and code benchmarks, MAD-OPD ranks first in overall average. On Qwen3 14B+8B→4B, it improves the agentic average by 4.39 points and the code average by 4.24 points over the stronger single-teacher OPD.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.