acceptodds
Under review as a conference paper at ICLR 2027

GMPROP: A LEARNED IMPORTANCE-SAMPLING PROPOSAL FOR PRF ATTENTION

Abstract

Softmax attention is powerful but expensive: its cost grows quadratically with sequence length. Positive random features (PRF) can approximate softmax attention and reduce its computational cost to linear time. To date, most methods used a standard normal distribution to sample random features. However, such an approach may poorly reflect the query and key distribution of a pretrained Transformer and therefore require many features for good accuracy. In this work we introduce GMProp. GMProp learns a sampling distribution from which the PRF are sampled. It represents the sampling distribution as a small mixture of Gaussians learned offline from the model’s query and key representations. At inference time, no fine-tuning or input-dependent optimization is required: the resulting method has the same linear-time attention computation as standard PRF attention and keeps its constant cost per generated token. Importance weights keep the kernel estimate unbiased for any learned proposal and any number of features. We provide theoretical justification for the design of GMProp and provide a thorough empirical investigation of the method. Across three pretrained language models, GMProp consistently improves over existing PRF baselines, and the learned proposals continue to perform well on a held-out corpus without being refit.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.