GIFT: Gradient Informed Filtering for Distillation Augmented Reasoning RL
Abstract
Training large language models for complex reasoning tasks relies on two complementary signals: sparse, outcome-grounded reinforcement learning (RL) for correctness, and dense, token-level distillation for guidance. However, complementarity in supervision does not guarantee compatibility in optimization; imitation-focused distillation can conflict with the outcome-driven RL objective, and existing hybrids lack a principled resolution. We propose **GIFT** (**G**radient-**I**nformed **F**il**T**ering for Distillation-Augmented Reasoning RL), a novel method that directly addresses the allocation problem of which teacher updates to apply. GIFT performs a token-level geometric analysis, decomposing the distillation gradient into a component **parallel** to the RL gradient, which measures first-order conflict and provides a **safety** filter, and a component **orthogonal** to it, which quantifies novel directional guidance and provides a measure of **value**. This decomposition yields a principled scoring rule to rank token-level guidance. We combine this with a progressive **coverage annealing** schedule that shrinks the budget of tokens receiving distillation, ensuring that the most conflicting updates are discarded first while the most complementary ones are retained at full strength. This facilitates a smooth transition from dense, geometrically filtered guidance to a pure, outcome-grounded objective at convergence. Extensive experiments on mathematical and logical reasoning benchmarks show that GIFT consistently outperforms strong baselines across multiple model scales, achieving relative gains of up to 9.5% on difficult mathematical reasoning tasks. *Code will be available upon publication.*
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.