acceptodds
Under review as a conference paper at ICLR 2027

The Convex Corridor: Learning Attention–Laplacian Mixtures Helps Pretraining

Abstract

Attention computes input-dependent weighted averages of token vectors through the softmax probability matrix . Recently proposed Laplacian heads instead use , where is the identity matrix, to compute the deviation of each token vector from its corresponding average. Although using both Laplacian and attention heads in transformers has been shown to improve performance on many tasks, how to best combine the two operators is still unclear. We address this problem by parameterizing each head's operator as , where are learnable parameters. This modification consistently leads to lower validation loss in causal language modelling pretraining across tested model scales. To understand why, we examine the evolution of and throughout training, and find that most heads beyond the first layer learn , an interval we call the convex corridor. When falls within this corridor, , so each head's operator becomes a scaled convex combination of attention and the negative Laplacian , with controlling the weights of the two operators and controlling the size of the head's output. This pattern emerges across different initializations and alternative parameterizations of the operator such as and , which also improve pretraining. We perform a series of controlled experiments to show that being in the convex corridor contributes to the improvements in pretraining: restricting increases loss, while fixing to be some value within the convex corridor already improves upon standard attention even with fixed to 1. Together, these findings suggest that attention should be viewed as one endpoint of the convex corridor, and that exploring the corridor’s interior is a promising direction for future transformer architectures.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.