WEIGHTED ON-POLICY DISTILLATION FOR MULTI- DOMAIN CAPABILITY INTEGRATION
Abstract
Multi-domain post-training seeks to combine specialist skills while preserving capabilities of the starting model. We propose Weighted On-Policy Distillation (WOPD), which evaluates specialists on shared verifiable training prompts and assigns domain-specific teacher weights in the OPD loss based on their success rates. The starting model receives a fixed weight as a reference teacher. All teachers supervise the same student-generated tokens through a weighted reverse-KL objective. In a five-domain study with Qwen3-4B, WOPD exceeds equal weighting on all five in-domain benchmarks by \(1.6\)–\(15.8\) percentage points under the same six-teacher pool. Compared with domain-routing Multi-Teacher On-Policy Distillation (MOPD), it improves the in-domain mean by \(1.1\) points and the out-of-domain mean by \(3.8\) points. WOPD matches the reference model on the out-of-domain math benchmark and exceeds it on the other four. Removing the reference teacher lowers every out-of-domain score by \(3.0\)–\(6.6\) points. At each student prefix, the weighted objective targets a geometric mixture of the teacher distributions, and reference supervision adds an explicit KL anchor to the reference model. Combining score-based specialist weights with reference supervision improves in-domain performance without out-of-domain losses relative to the reference model.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.