Sequential Stability Bounds for Order-Dependent In-Context Learning
Abstract
In-context learning (ICL) accuracy swings by up to 39 percentage points under demonstration reordering, yet every existing theoretical framework treats the prompt as an unordered set and therefore certifies one number for all orderings. We formalize causal attention masking as the source of this asymmetry: tracking a single-example replacement through the causal chain gives a closed-form position-dependent amplification factor , strictly decreasing in slot for and exactly flat for , which couples with a per-source sensitivity score to define *sequential stability coefficients* —the first bounded-differences constant for ICL that varies with the permutation. Three theorems follow: an order-dependent generalization bound, a rearrangement-based certification of low-sensitivity-first ordering with an separation, and a non-vacuity criterion in that is independent of embedding dimension; a spherical-concentration certificate resolves distinct sources at cost . A hierarchy of diagnostic experiments matches the theory: produces spread below across 120 permutations, matched trained and untrained networks are equivalent under normalized attention (TOST , )—so ordering effects are architectural, not learned—and four 7–8B instruction-tuned LLMs show 8–39% accuracy spread on AG News and RTE. The multiplicative decomposition is recovered from the empirical gap matrix: the position factor tracks closed-form at –, an intraclass ceiling – quantifies the per-example share, and a learned score saturates of that ceiling at —identifying the residual as the framework's predicted many-body component.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.