Taylor-Calibrate: Principled Initialization for Hybrid Linear Attention Distillation
Abstract
Hybrid linear attention reduces the quadratic cost and KV-cache growth of softmax attention, but converting a pretrained Transformer remains brittle. Projection transfer copies weights yet leaves the student's new recurrent dynamics—memory decay, write strength, value scale, and output gating—unspecified. We propose Taylor-Calibrate, a lightweight teacher-conditioned initialization for these unmatched dynamics. An exact online form of softmax attention exposes a delta update; a Taylor and online-regression approximation then identifies the corresponding recurrent quantities. We estimate their operating point from teacher attention distance, concentration, and output scale, followed by brief layer-local alignment. Across four teachers and three retained-layer policies, Taylor-Calibrate produces stronger zero-shot students. On five single-run matched-PPL recovery comparisons spanning three teachers, conservative lower-bound Stage-2 token savings range from 4.9 to 9.2 relative to projection-copy initialization. Held-out representation and KL measurements confirm that the gain reflects closer teacher behavior, not perplexity alone. Matched three-seed experiments further transfer the framework to GLA and KDA. Separately, five-layer KDA conversion with a 64-GPU-sharded 975B Inkling teacher reaches PPL 6.498 versus 6.454 and 90.2% teacher top-1 agreement, validating scale integration. The complete 1.5B calibration costs 0.024 H100-hours.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.