Information Limits and Output Geometry for Quantized Attention
Abstract
A certificate can reject a quantized attention output even when its actual error is small. We characterize how retained probability and value information determines the tightest attainable error bound. Our exact scalar output envelope combines range, mean, and variance constraints: support bounds clip the classical variance-only optimizer, and two-point laws attain every regime. Joint constraints can certify budgets that separate moment bounds cannot. A distinct construction gives a separation from the output conversion of the optimal total-variation (TV) bound. We also derive attainable TV envelopes for selected exact probabilities and lower bounds, with saturation and stopping criteria. Outward moments and a sharp correction for perturbed-distribution value error connect the envelope to quantized attention. On a fixed GPT-2 replay, an exact CPU reference certifies 357 of 1,008 case/budget pairs, versus 249 for midpoint-centered TV and 352 for coordinatewise mixing of variance-only and range/mean bounds. Independent numerical enclosures validate every certified pair. These comparisons distinguish the benefit of value information from the additional gain of joint constraints.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.