acceptodds
Under review as a conference paper at ICLR 2027

Nexus: Higher-Order Attention Mechanisms in Transformers

Abstract

Transformers have achieved broad success across language and vision, using self-attention to model dependencies among input elements. Yet the relevance of one position to another often depends on evidence distributed across a wider context. Standard attention forms each matching score from a query and a key projected at the current layer, whose inputs may already carry context from earlier layers. Context-aware transformations and multi-token interactions enrich this process, motivating further study of how context can enter the matching operation itself. We introduce Nexus, a higher-order attention mechanism that refines queries and keys through separate causal attention branches before their final comparison. The construction extends recursively while reusing the original projections. Our analysis shows how other prefix tokens can influence comparisons within a single attention head. Experiments across four from-scratch model scales yield higher six-task averages than standard attention, and adapting pretrained Qwen2.5 models with Nexus improves reasoning-task averages after supervised fine-tuning. These results support recursive query/key refinement as a design for context-dependent matching in Transformer decoders.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.