acceptodds
Under review as a conference paper at ICLR 2027

Dissociated Recurrent Clustering

Abstract

The current status quo for large-scale recurrent models is to restrict the recurence to linear operations, which are associative and can be implemented efficiently. We depart from this key constraint with the Dissociated Recurrent Clustering (DRC), whose output is a standard attention over a recurrent state composed of a set of weighted key/value centroids that are iteratively fused to keep their number fixed. To make this process parallel during training, we introduce the “dissociation” of the computation: the centroids to fuse are computed on very large batches without autograd and, once they are fixed, the forward and the backward can be formulated as parallel scans. Experiments with hybrid 1.5B models trained on sequences of 32k tokens and evaluated on standard benchmarks, without hyper-parameter tuning, show that using DRC with 1k centroids instead of Gated DeltaNet with an equivalent recurrent state size reduces the gap to using a standard attention layer by a factor 2 in term of training loss and 3.5 in validation loss on source code, known to require long context operators.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.