acceptodds
Under review as a conference paper at ICLR 2027

Learning Forgetting Rates on a Half-Life Grid

Abstract

Recurrent language models such as GatedDeltaNet compress context into a fixed-size state and use a learned forget gate to control its decay. In the GatedDeltaNet models we train, the learned gates favor short retention times: the median head half-life is only 9–16 tokens. We propose Mixture Half-Life Decay Gates (MHLG), in which a token-conditioned router sets each head's decay rate to a weighted average of the rates corresponding to a small, ordered grid of learned half-lives. Each head applies a single scalar decay to a single state, so the state size and update rule are unchanged, while the resulting half-life stays between the shortest and longest learned half-lives and the router can retain weight on slow timescales. In a study of capacity-matched 134M-parameter models trained on 262M tokens, MHLG lowers WikiText-103 test perplexity by 11.5 points, with all five paired seeds favoring it, and by 11.8 points on the full test set. The WikiText-103 improvement is consistent across all evaluated seeds at 400M parameters, at about 1B training tokens, and on a second training corpus. Over three seeds, MHLG reduces perplexity by 9–12 points relative to a wider continuous gate, a trained continuous router, and a control initialized on the same half-lives, and, per layer, its router places 33–85% of its weight on the slowest half-life. A variant that fixes half of the router weight on the slowest half-life raises the gain to 14.5 points. These controls point to a bounded forgetting rate with a maintained slow end as the source of the improvement, and to a learned half-life grid as an inspectable parameterization of this constraint.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.