acceptodds
Under review as a conference paper at ICLR 2027

From Frequency to Function: Component-Specific RoPE Extrapolation for Long-Context Language Models

Abstract

Rotary Position Embedding (RoPE)-based language models form position-sensitive attention by superimposing contributions from dimension pairs associated with different rotation frequencies. Long-context extrapolation therefore requires rescaling rotation frequencies. However, common frequency-based methods partition dimensions by the relative ordering of their associated RoPE frequencies and scale attention logits with a single global factor. This design overlooks the distinct functional roles that dimension pairs acquire during pretraining. We group dimension pairs into three classes: positional pairs have concentrated pre-RoPE activation phases and produce stable relative-position periodicity, semantic pairs are dominated by input content variation, and sink pairs make exceptionally large positive contributions to sink-token attention logits. Based on this decomposition, we propose FARoPE (Function-Aware RoPE), a training-free RoPE extrapolation method that first identifies the three pair classes and then applies component-specific transformations. Specifically, FARoPE applies a shared frequency scale to positional pairs to prevent spurious shifts in position-sensitive attention patterns. Sink pairs receive the same frequency reduction to maintain sink alignment, while semantic pairs retain the YaRN ramp. FARoPE separately calibrates the additional attention-logit gain for sink pairs without masking sink tokens. Our evaluation covers two model families and three long-context benchmarks with context windows up to 128K.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.