When Is an Attention Sink Robust? A Convex-Geometric Theory
Abstract
Attention sinks are token positions that receive disproportionately large attention across queries. Yet it remains unclear when a sink is robust, whether it is arbitrarily amplifiable, and how key perturbations and context growth change the required query norm. We answer these questions for a fixed, pretrained attention head under deterministic product uncertainty with compact convex key sets. We prove that the queries at which the candidate token is a robust attention sink at a prescribed confidence level form a closed convex region. Our first main theorem gives an exact entropy–distance formula for the minimum query norm needed to make the candidate token a robust attention sink. The distance term measures the separation between the sink and weighted combinations of competing-token keys, while the entropy term captures their collective softmax contribution. The theorem also determines exactly when the robust attention weight of candidate token can be amplified arbitrarily close to one. Our second main theorem determines how the minimum query norm needed to preserve the candidate token as a robust attention sink changes under key perturbations and context growth. Let be the key-error radius, the number of added competing tokens, and the nominal separation from the sink key to the aggregate competing-key set. For fixed , the required norm grows as ; for fixed , it grows as as approaches , and no finite query norm suffices when . Together, our results distinguish whether a robust sink is geometrically possible from how large a query norm is needed to realize it, providing a principled basis for auditing sink preservation under key errors and long-context inference.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.