Delayed Positional Adjustment for Context Extension
Abstract
Extending a pretrained language model to longer inputs changes two things at once: the training sequences grow longer, and the rotary position embedding (RoPE) frequencies are rescaled to a target such as position interpolation (PI) or YaRN, usually from the first update. We ask when to introduce the target frequencies. We introduce delayed positional adjustment (DEPA), which trains at the new length with the original frequencies, switches once to the target after 90% of the continuation updates, and adapts for the remaining 10%. We compare DEPA with static adjustment, which applies the target from the first update, using pairs matched on the pretrained model, final frequencies, token budget, and optimizer schedule. On 137M-parameter models extended from 1K to 4K tokens on PG19, DEPA lowers perplexity (PPL) beyond the trained window in every tested pair. Under our custom YaRN-family target (a YaRN variant with our own frequency mask), whose static runs already extrapolate well, five constant-learning-rate pairs show a reduction of about 1.9 in 16K-token PPL at a 0.03-PPL cost at 4K. Under PI and published YaRN, DEPA lowers 16K PPL from about 105–190 to below 32. The gain persists on a second pretrained base, on WikiText-103, for a 20-layer model, and at 19M parameters. When fixed target tokens are scored with progressively longer inputs, DEPA reduces added-context harm by 78–80% on PG19. Release timing is a consequential, easily controlled choice that context-extension recipes should set and report.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.