acceptodds
Under review as a conference paper at ICLR 2027

TITRATE: Adapting Self-Distillation Weight and Skill Refresh to Training Dynamics

Abstract

Skill-augmented On-Policy Self-distillation trains large language model (LLM) agents by combining Reinforcement Learning (RL) with guidance from a teacher that shares the student policy's parameters but receives hindsight-derived skills as privileged context. However, fixed skill-refresh rules and distillation weights leave skill evolution poorly coordinated with policy learning. We introduce , a unified framework that coordinates these processes through: (1) a that uses an analyzer to consolidate successful and failed rollout evidence at task or class granularity and supports incremental, version-compatible updates; (2) an controller that uses rollout evidence and confidence-based ranking to prioritize skill refreshes under analyzer-call budgets; and (3) a that combines progress-aware cosine annealing with bounded, conflict-aware refinement to balance skill-informed guidance and reward-driven learning. Together, they form a closed co-evolutionary loop: accumulated experience informs skill refinement, skill-informed self-distillation guides policy updates, and the updated policy, in turn, generates new experience. On ALFWorld, WebShop, and SearchQA with three Qwen models (1.7B–7B), TITRATE achieves the highest mean task performance in eight of nine settings, while reducing analyzer calls by 88.8–99.4% and analyzer token consumption by 8.1–56.7% relative to per-trajectory extraction. On class-level SearchQA with QWen2.5-3B, TITRATE ranks second while enabling mixed-outcome updates that restrictive baselines suppress despite unresolved failures. Ablations further support the contributions of evidence-gated refresh and adaptive weighting. Our results demonstrate that coordinating skill evolution and distillation strength with policy learning can achieve a better performance–cost trade-off, supporting TITRATE as an effective solution for self-distillation RL training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.