acceptodds
Under review as a conference paper at ICLR 2027

Scaling Laws For Attention-To-Recurrence Distillation

Abstract

An efficient way of reducing the computational cost of pretrained large language models is to generate hybrids by replacing some of the traditional attention layers, whose complexity scales quadratically with context length, with recurrent alternatives such as state-space models, which scale linearly with context length. Approaches such as MOHAWK allow distilling hybrids under limited token budgets compared to output-level Knowledge Distillation (KD), while retaining model capabilities. Yet, they lack the theoretical foundation needed to guide practitioners on how many full-attention layers can be replaced and at what cost. Here, we aim to close this gap. Using arguments from spectral theory, we argue for a conditional lower bound on the Kullback–Leibler (KL) divergence loss between hybrid and non-hybrid models in a restricted componentwise setting, connecting it to the attention matrix approximation error in MOHAWK’s matrix-orientation phase. We also study a scaling law for the context-dependent KL loss based on the ratio of replaced layers, their expressivity and the Phase 2 token budget. Finally, we verify experimentally that the KL divergence loss often exhibits power-law decay within a fixed conversion setup, while finding that its late-budget shape does not transfer cleanly across model sizes. Overall, our work sets the foundation for a theory that characterises the practical utility of distilling hybrids from pretrained models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.