acceptodds
Under review as a conference paper at ICLR 2027

MINKOWSKI ATTENTION

Abstract

This paper introduces Minkowski Attention, a novel formulation for attention and self-attention layers through a generalization of Minkowski spacetime. Attention layers must reason over two distinct representations per token, for content and position. The canonical fix, where content and position embeddings are summed before computing attention, entangles these two signals into cross terms that don't cleanly separate, hence a growing body of work advocates for decoupled attention. Decoupled attention however requires computing separate attention matrices for each pair of representations, inducing considerable compute and memory cost per layer. In this work, we show that with a Minkowski spacetime perspective, it becomes possible to recover a single attention score from decoupled representations. We propose to view token and positional embeddings as multi-dimensional space and time vectors, where token pairs are scored with a Minkowski inner product instead of a Euclidean dot product. The signed structure of Minkowski spacetime keeps content and positions contributions separable under one scalar. We generalize the (3+1)D Minkowski metric to arbitrary (p+q)-dimensional signatures and to the corresponding generalized Lorentz boosts. Our Minkowski transformer consistently outperforms coupled and decoupled transformers across various tasks and modalities, demonstrating strong attention diversity over attention heads and robustness to non-discriminative tokens.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.