acceptodds
Under review as a conference paper at ICLR 2027

Computation Persists, Routing Degrades: A Mechanistic Analysis of Transformer Length Extrapolation

Abstract

Transformer models trained on algorithmic tasks often fail to extrapolate to sequence lengths beyond those observed during training. This behavioral degradation is frequently interpreted as a fundamental collapse in reasoning capacity, but we investigate whether this failure is mechanistically decomposable. Using dynamic programming (DP) tasks as case studies, we analyze the internal execution of RoPE-based Transformer model organisms. We find that length extrapolation failure is heterogeneous across learned mechanisms. Specifically, in the studied Longest Increasing Subsequence (LIS) models, internal representations associated with DP state construction remain substantially decodable beyond the training horizon, while downstream positional routing mechanisms degrade sharply. Crucially, through out-of-distribution (OOD) activation patching, we demonstrate that this unmodified OOD state representation remains functionally usable: correcting downstream routing alone recovers substantial behavior relative to the OOD baseline, and upstream restoration provides no additional benefit under this intervention. Furthermore, evaluating sequences at a fixed length reveals that dependency distance is a key correlate of representational degradation. In contrast, on the Longest Common Subsequence (LCS) task, where state construction itself requires global cross-sequence interactions, both intermediate representations degrade alongside behavioral performance under extrapolation. Layer-specific positional interventions and Randomized Positional Training further implicate positional geometry as a major constraint. Together, these results suggest that algorithmic extrapolation failure depends on the positional interaction requirements of individual learned mechanisms, rather than reflecting a uniform loss of algorithmic information.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.