Attention Is Not All You Need Without Structure: Neuro-Symbolic Multi-Head Attention
Abstract
Transformer architectures rely on Multi-Head Attention (MHA) as the central mechanism for modelling interactions between tokens in sequence data. Although this is effective for capturing statistical dependencies, the standard MHA lacks explicit structural inductive bias for relational reasoning, which limits systematic generalisation for tasks requiring multi-step inference. Prior neuro-symbolic methods attempt to fix this limitation through external reasoning modules or discrete symbolic pipelines, which can complicate integration with modern transformer architectures and break end-to-end differentiability. In this work, we introduce a role-conditioned extension of MHA that injects structured relational inductive bias directly into the attention computation while preserving full end-to-end differentiability along with transformer compatibility as a drop-in replacement. The proposed mechanism augments standard attention with a parallel role-conditioned projection pathway based on learned role embeddings, combined through a learnable fusion gating mechanism. Additionally, a role-relational attention bias enables symbolic attention weights to depend on interactions between inferred token roles, allowing relational structure to influence attention scores directly. We evaluate the proposed architecture on language modelling, relational reasoning, and compositional generalisation benchmarks, including WikiText-2, LAMBADA, CLUTRR, and SCAN, using multi-seed experiments with paired statistical testing. Results demonstrate consistent reductions in perplexity on language modelling benchmarks, substantial improvements on relational reasoning tasks, and improved probabilistic modelling on compositional tasks. Ablation studies confirm that the observed gains arise from role-conditioned structural bias rather than architectural duplication alone. These findings demonstrate that structured inductive bias can be embedded directly within attention mechanisms, enhancing Multi-Head Attention while preserving end-to-end optimisation and transformer-native design.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.