LOTFormer: Doubly Stochastic Linear Attention via Low Rank Optimal Transport
Abstract
Standard softmax attention is row stochastic and scales quadratically with sequence length. Doubly stochastic attention also controls aggregate key participation, but computing it is typically expensive. We introduce LOTFormer, an efficient attention mechanism designed for bidirectional Transformers, where tokens act as both information sources and receivers. LOTFormer routes interactions through a learned low rank pivot measure and composes two entropic transport couplings. The resulting ideal attention operator is doubly stochastic, has rank at most , and applies to values in time without forming the full pairwise matrix. We evaluate the method both as a replacement for attention in pretrained embedding models and as an attention mechanism trained in vision Transformers. Across embedding models from 1.5B to 7B parameters, LOTFormer preserves to of the original softmax retrieval quality. On four natural LongEmbed tasks, gte Qwen2 7B retains of its original nDCG@10. In a separate throughput stress test at 128k tokens, it achieves a speedup over softmax. LOTFormer also performs competitively across multiple vision backbones and under limited supervision. These results demonstrate a practical route to efficient bidirectional attention with doubly stochastic structure. Code is available at https://github.com/anonymous13761234/LOTFormer.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.