Natural-Gradient Calibration of Multi-Turn On-Policy Distillation
Abstract
On-policy distillation trains a language model by matching a teacher on the student's own rollouts, and the dense token-level signal complements the sparse terminal reward of reinforcement learning. Rollouts of multi-turn agents divide into turns at tool-call boundaries. Recent methods reweight the teacher-matching loss turn by turn from a gap between teacher and student probabilities, from the terminal reward, or from a depth schedule, without estimating how the teacher supervision at a turn contributes to task success. We view on-policy distillation as a policy-gradient step whose advantage is a proxy, the weighted distillation signal, so a turn-level weight scales that proxy toward the true advantage. The weight that fits the proxy best under the student's action distribution is the covariance of the signal with the return divided by the variance of the signal, which is the projection coefficient of the reward gradient onto the distillation direction in the Fisher metric. The projection splits the policy gradient into a calibrated distillation direction and a residual identified by the return alone, so calibrated distillation is the part of reinforcement learning that the teacher supplies. The nonnegative parts of the per-turn projections are also the assignment that gains the most reward in the second-order KL geometry of one update. We instantiate the calibration as *Turn-Level Advantage Calibration in the Trust Region* (**TACT**). The loss already computes the divergence at each position, and adding it to the token log-ratio makes one rollout an unbiased sample of both the covariance and the variance of every turn it visits. TACT therefore estimates the weights from the training batch, pooled over classes of turns, and adds a reinforcement step on the residual in the classes where the signal explains little of the return. Experiments on ALFWorld, ScienceWorld, WebShop, Search-QA, and -bench, on the same rollouts for every method, show that TACT exceeds the strongest existing rule on every environment, by 4.9 points on average over the rule that uses the return, at the same cost as on-policy distillation. The code is available at [TACT Codebase](https://anonymous.4open.science/r/TACT_ICLR-BD0D/README.md).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.