acceptodds
Under review as a conference paper at ICLR 2027

Virtual Tokens, Real Supervision: Task-Aware Extrapolation for Vision RoPE

Abstract

To improve generalization beyond the training scale, recent methods introduce Vision RoPE into vision models. However, at substantially larger spatial or spatiotemporal resolutions, existing extrapolation methods typically optimize from only one of two signals: source-scale task supervision or target-grid geometry. Source-supervised methods do not observe the target scale, whereas target-aware post-hoc corrections usually rely on geometric or statistical heuristics without target-conditioned task supervision. Their inability to use both signals simultaneously limits the effectiveness of extrapolation optimization, especially at large resolutions. We present Virtual Target Calibration, a source-only post-hoc method that constructs a target-conditioned surrogate from source-scale inputs. VTC stretches RoPE coordinates and assigns analytic multiplicities to visual keys to model target-scale phase changes and enlarged softmax competition. Source-task Fisher weights then guide the joint calibration of a block-unitary frequency mixer and a positive attention inverse-temperature factor. Calibration is performed at a small set of target anchors and interpolated to unseen resolutions while the backbone remains frozen; classification calibration takes about five minutes on RTX 4090 hardware, with at most 200 updates per anchor. Across image classification, high-resolution image generation, and long-video generation, VTC improves extrapolation over native inference and is competitive with or better than heuristic post-hoc corrections, without target-scale inputs or labels.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.