SV-ROPE: TOKEN PROFILES OF SPATIAL VARIATION FOR REMOTE SENSING FOUNDATION MODELS
Abstract
Remote-sensing foundation models (RSFMs) acquire transferable spatial priors from large-scale Earth observation (EO) imagery, which are commonly inherited during downstream parameter-efficient fine-tuning (PEFT). Yet, for Transformer-based RSFMs, the spatial meaning of the same index-space token relation can vary substantially across observations, leaving inherited priors poorly matched to target-specific spatial variation. This raises a central question: how can inherited priors be made responsive to target-specific spatial variation within a PEFT adaptation framework? We therefore introduce a spatial variation-aware view of RSFM adaptation and propose SV-RoPE (Spatial-Variation-Aware Rotary Position Embedding), a PEFT method that associates each visual token with a spatial-variation profile (SV-Profile). Each profile summarizes how the token's feature agreement with neighboring tokens changes across increasing spatial extents. Leveraging the coupled displacement–frequency structure of RoPE, SV-RoPE encodes spatial variation directly into a unified rotary phase, jointly parameterized by Relation Coordinate Calibration (RCC) for initial ground-extent-calibrated displacement and Relation Frequency Calibration (RFC) for SV-Profile-conditioned frequency modulation. SV-RoPE is realized through a lightweight attention-parallel branch whose output is added to the attention of the frozen backbone. Across diverse RSFMs, sensors, resolutions, and downstream tasks, SV-RoPE consistently improves adaptation performance, demonstrating the value of explicitly modeling spatial variation in downstream EO adaptation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.