acceptodds
Under review as a conference paper at ICLR 2027

FasterWeight-MLP Attention: Sequence Modeling from a Neural-Network-Native Perspective

Abstract

A causal attention layer computes its scores with a network whose weights the context instantiates: a softmax head is a two-layer network whose hidden units and readout come from the prefix. We take that network as the design object and ask two questions: what may it read, and how much learned capacity does it carry? A query position and a key position split the causal history into three regions: the prefix the query sees, the prefix the key sees, and the interval between them. FasterWeight-MLP Attention (FW-MLP) reads each region with a small learned network, a fast MLP, and multiplies the three outputs into one score. Reading the interval by subtracting continuous prefix statistics leaves one exact form, a running additive sum, and with it FW-MLP evaluates the score exactly by a recurrence linear in sequence length. Linear attention, retention, and gated decays are exact special cases of the resulting score class, and delta-rule transports and the continuous scores on compact domains lie in its uniform closure. Under a matched attention-side parameter budget, FW-MLP attains the best mean rank among 12 mixers on 26 tasks and the best validation loss in a 123M-parameter nanoGPT setting. On a retrieval task gated by an interval statistic it reaches 34.3% accuracy, against 20.4% for its endpoint-only configuration, and it holds the best average downstream rank at both 340M and 1.3B parameters.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.