Exact-Solution Volume and Length Generalization in Transformers
Abstract
Research on transformer expressivity shows whether a transformer is capable of solving a given task, but gives little indication of whether the solution, if learned, is generalizable to longer input lengths. We study this question through *normalized exact-solution volume (NESV)*: the fraction of a bounded parameter region that achieves an exact solution on every input of length . For fixed-width, single-layer transformers with -scaled attention, we establish asymptotic bounds on NESV for four tasks: FIRST (), MAJORITY (), INDEX (), and PARITY (). These results are consistent with previous empirical results: the faster the NESV decays with input length, the harder it is to length-generalize on that task. Looking deeper into INDEX, our volume analysis reveals two error sources that grow with . Consequently, we study a transformer model that would structurally eliminate one of the terms, theoretically improving the NESV bound to , and empirically achieving 85% accuracy when tested at 10 times the training length, compared with the 60% accuracy of the original model. We conclude that volume analysis may be a useful approach to identify concrete sources of length sensitivity and thus provide insights into task-specific model refinements.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.