When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation
Abstract
Teacher top-K logits can preserve nearly all teacher probability mass while omitting corrections needed for a student's behavioral decisions. We study this gap in routed, two-teacher tool-use distillation: a tool teacher reinforces an entry coordinate that the response teacher usually omits even when the student overestimates it. In Qwen3.5-9B, response-teacher top-32 retains 99.99% of its mass yet contains the tool-entry token on only 0.4% of 500 response prompts. A frozen full-vocabulary audit shows that restoring this coordinate with its exact teacher logit recovers full-vocabulary descent on that coordinate. Forced-entry replay and matched training connect the omission to behavior: first-position restoration relocates calls to later tokens, whereas all-position restoration reduces full-generation over-calling from (14.2 +/- 2.1)% to (3.7 +/- 0.5)% while lowering required-call recall by 12.4 points. Teacher-probability-matched and output-logit-gradient-rerouted non-tool controls produce smaller shifts and do not reproduce exact restoration. We then compare support construction, loss shaping, and validation-tuned decoding bias. The comparison separates changes in the call threshold from modest gains in discrimination and exposes the access and capability costs of reducing calls. Llama-3.1-8B reproduces the directional support asymmetry under native JSON. Together, these results identify decision-critical support omission as a causal contributor in the primary Qwen setting and show why teacher-mass coverage alone cannot certify behavioral fidelity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.