Delay-SSD: Reading Past the Decay with Lagged State Updates
Abstract
Structured state space models such as Mamba-2 forget under one scalar decay per head, so a token's influence is fixed by its distance in the scan. A dependency at a fixed lag , such as the pixel above in a flattened image or the period of a seasonal series, must therefore be carried through decay steps and competes for state with every token in between. We introduce , which writes the input from steps earlier directly into the SSD state update, so that the lagged token is read at decay distance zero. All lags share one state and one decay, so the state size is unchanged and Mamba-2's chunked scan still applies, because the sequence transformation stays semiseparable of order outside a band of width . We show that neither more state nor another scan order buys this reach. With input-independent parameters, scalar-decay heads cannot isolate a lag beyond at any state size, input-dependent heads reading independent tokens are bound the same way on average, and no input-independent linear recurrence with fewer than states approximates a -step delay, whereas a single lag does both. On synthetic routing tasks, a scan order changes which offsets one SSD layer reaches but not how many, while one layer also reaches the rows above and below. With our TileLang kernels, two lags cost – the device time of the four-scan SSD mixer, compared with about for a second layer or a larger state. As a vision backbone, matches Mamba2D on ImageNet-1k ( top-1 at 50M parameters) and COCO, and leads it on ADE20K segmentation at every size (by , and mIoU). In forecasting, with only the mixer changed, lags lower the error on each of five traffic datasets, by relative to a no-lag control and by relative to Mamba-2.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.