Attention as Active Stochastic Control: A Unifying Theory and Architectural Framework
Abstract
For nearly a decade, the Transformer architecture has served as the empirical workhorse of modern artificial intelligence. Despite its ubiquitous deployment, its core mechanism, softmax self-attention, is primarily understood as a heuristic pairwise sequence mixer optimized via gradient descent. Consequently, pervasive failure modes such as prompt fragility, factual hallucinations, representation collapse (over-smoothing), and attention sinks have been addressed largely through ad-hoc engineering heuristics. In this work, we establish a rigorous mathematical unification between Transformer Attention and Active Stochastic Control. Specifically, we demonstrate that along the depth axis, self-attention operates as an *active observer* executing dual control: the query functions as an exploratory control action steering sensor admittance to minimize epistemic entropy, the key matrix defines a directional channel alignment metric, and the softmax operator arises as the exact analytical Bayesian update of a Wonham filter for categorical point-process emissions. Furthermore, viewing residual stacks as continuous-time Hamiltonian interacting particle systems exposes a symplectic leapfrog structure enabling activation memory training. Leveraging this framework, we introduce **Closed-Loop Iterated Active Attention (IAA)**, which resolves multi-hop relational reasoning within a single physical layer, reuses a static key-value observation manifold, and terminates computation dynamically via Kalman innovation norm decay. Finally, we translate foundational principles from robust control (), balanced truncation (Hankel singular values), and Bode's sensitivity integral to provide principled explanations and actionable solutions for transformer jailbreaks, KV-cache compression, and attention sinks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.