acceptodds
Under review as a conference paper at ICLR 2027

CyFA: Linear Sequence Modeling with Relative-Time-Partitioned Memory

Abstract

Linear recurrent models compress history into a fixed-size state, forcing the same finite memory to accommodate an ever-growing sequence of writes. Existing updates handle this pressure through superposition, forgetting, or replacement, which can make earlier associations difficult to recover. We introduce CyFA (Cyclic Flow Attention), which instead reuses memory by cyclically transporting the surviving history as new information arrives. Earlier writes move away from the age-zero end of the memory while new key–value pairs enter there, producing a relative-time-partitioned state in which multiple historical associations can coexist across the cycle. An exact change of coordinates reduces this moving memory to two scalar-decay linear attention recurrences, enabling efficient chunk-wise training. Across 400M–1.4B pretraining experiments, CyFA consistently outperforms linear attention models on recall-intensive tasks under matched recurrent state sizes. At 400M, it reaches on FDA versus for KDA, while maintaining strong language modeling and higher training throughput than models with fine-grained gating or Delta Rule updates.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.