acceptodds
Under review as a conference paper at ICLR 2027

When Does a RoPE Extrapolation Repair Help?

Abstract

Context-extension repairs for rotary position embeddings (RoPE) can rescue a model whose predictive performance collapses on long documents, yet harm one that already handles long inputs, so practitioners must know in advance which case applies. We study a diagnostic that predicts when one low-budget repair helps. The repair retains RoPE, adds to each attention head a distance penalty scaled to how far back that head attends, and trains for 30 updates. The diagnostic is the matched no-penalty continuation's ratio of perplexity after a document's first 2,048 tokens (late-token perplexity) to perplexity on those initial tokens. At a ratio of at least 2, a fixed rule fires and predicts at least 5% lower late-token perplexity than that continuation; otherwise, it abstains. On 30 checkpoint–length cells from ten public checkpoints, including five named before evaluation, the rule fires on 20 and all 20 improve; it abstains on exactly the 10 cells without a reduction, including native long-context Qwen3 checkpoints, where the repair raises perplexity. On three independently trained 124M parents, the repair lowers 16K late-token perplexity from 168–182 to about 25 in 74 GPU-seconds, at a cost of about 0.1 in in-window (PG19) perplexity; measured slopes beat three calibration-free schedules on late-token perplexity for every parent. In a key-swap control, a correct key just before the question raises the repaired models' answer log-likelihood 7.3–8.2 nats more than a wrong key. On Pythia-1b, the repair has the lowest in-window perplexity of the compared extensions. A single paired run thus makes the decision to apply this repair testable.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.