acceptodds
Under review as a conference paper at ICLR 2027

Learn from Your Better Self: On-Policy Recovery for Sparse Language Models

Abstract

Sparse attention reduces inference cost but can induce *sparse reasoning drift*, where early token deviations compound through autoregressive feedback, degrading reasoning and prolonging generation. We introduce sparse on‑policy distillation (**SparseOPD**), which uses the original checkpoint's dense execution as a privileged teacher to supervise next‑token distributions on trajectories generated by the evolving sparse student. Across Qwen3‑4B, Qwen3‑14B, and Ministral‑3‑14B, SparseOPD improves mathematical reasoning and code generation over cross‑entropy recovery and fixed‑trajectory distillation. On AIME 2026, it nearly restores dense accuracy on Qwen3‑4B and exceeds it on Qwen3‑14B. It also curbs excessive generation and retains a request‑level speedup over dense inference on Qwen3‑4B HMMT. Our analyses localize on‑policy gains to later sparse states and reveal that recovery signals concentrate at low‑probability token positions. This motivates **LowP‑40**, which supervises only the lowest‑probability of tokens. On Qwen3‑4B, LowP‑40 maintains recovery quality comparable to full‑token supervision while accelerating end‑to‑end training by .

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.