acceptodds
Under review as a conference paper at ICLR 2027

Training-Free Long-Context Extension by Warping Rotary Relative Distance

Abstract

Rotary position embeddings (ROPE) are the default positional mechanism in modern language models, but training-free context extension remains brittle: global position or frequency rescaling can bring distant tokens into the trained range while distorting the local geometry the model learned during pretraining. We introduce WARP (Warped Rotary Positions), a training-free method that keeps the pretrained rotary spectrum fixed and instead applies a monotone nonlinear warp to the relative query–key distance used inside rotary attention. WARP is exactly ROPE at native context length, approximately preserves local distances, and compresses only the far field into the model’s trained phase budget. Because the warp is a nonlinear function of the pairwise distance i − j, it cannot be represented by a standard per-token position cache, which makes WARP a distinct class of rotary extension from position or frequency rescaling. The evaluated configuration adds a small frequency, temperature, and density calibration and applies the warp through a block-anchored approximation. In ten-seed exact-retrieval stress tests, WARP reaches 50/50 with zero observed variance on Qwen2.5-3B-Instruct and Mistral-7B-v0.3 at both 32k and 64k and on Qwen2.5-0.5B-Instruct at 64k, and reaches 49.20/50 on Llama-2-13B at 16k; Llama-2-13B at 32k remains weak. On official RULER v1 at 64k with Qwen2.5-3B-Instruct, WARP achieves 70.82% macro accuracy, compared with 62.92% for unmodified ROPE, 71.22% for NTK, 73.31% for YaRN, and 73.44% for MrRoPE-Pro. WARP is therefore competitive on broad RULER rather than dominant; its advantage is concentrated in retrieval-heavy behavior, where it is best or tied-best among the strongest methods on several RULER retrieval and QA components. The hardest setting we observe is Llama-2-7B exact retrieval at 32k, where base WARP and every baseline score 0/50; an exploratory variant that adds head routing and a small trained readout adapter recovers up to 31/50. These results suggest that relative-distance warping is a promising, model-adaptable route to long-context extension without finetuning, while exposing architecture-specific failure modes that we report explicitly.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.