acceptodds
Under review as a conference paper at ICLR 2027

A Continuous Space of Attention Projection Geometries at Initialization

Abstract

Standard Transformer initializers set how large the query and key weights are, but not how those two matrices relate to each other. Attention scores, however, depend on their product. We treat that missing relation as a continuous family: a shared component, a residual orthogonal to the query row space, and an independent residual, all with the same expected weight energy. Closed-form finite-width formulas give the mean and variance of every attention score under this family. Monte Carlo checks match the formulas, and a sweep of the family at a fixed width and context produces several distinct initial attention patterns which are nearly uniform, sharply peaked, self-focused, and a mid-entropy band in between. The same coefficient values do not keep those patterns when width and context change. A one-block training comparison that first looks like a geometry effect is explained by scale: mid-entropy and diffuse points sit at small scale and train similarly, while concentrated and self-focused points were copied in at much larger scale and train worse. Joint querykey geometry is a real finite-width degree of freedom. The map we report describes one setting; it is not yet a portable initialization recipe.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.