Beyond Training Lengths: Conditions for Generalization
Abstract
A key challenge in learning-to-reason tasks is that models trained on shorter sequences often fail to handle longer instances at test time. This phenomenon, known as length generalization (LG), requires extrapolation beyond the training length, whereas most learning paradigms are designed for interpolation under the i.i.d. assumption. For this reason, LG is also referred to as length extrapolation. Despite substantial empirical progress in recent years, LG remains largely unresolved. While prior theoretical work has examined learning to reason, relatively little attention has been paid to reasoning under LG. In this paper, we study necessary and sufficient conditions for achieving LG in a manner that is independent of specific learning algorithms or model architectures. This perspective is important, as it avoids limiting the analysis to the capabilities and constraints of existing learning algorithms. Finally, we analyze two empirically successful LG approaches based on the proposed necessary and sufficient conditions, providing initial evidence for the relevance of our theoretical framework.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.