acceptodds
Under review as a conference paper at ICLR 2027

Foveal Re-Encoding for Frozen Backbones

Abstract

Human vision is foveated, resolving detail only around the point of gaze and moving it wherever the task needs acuity, whereas a vision foundation model encodes one global view into a low-resolution grid and can look closer only by pushing a crop through the backbone again. Feature upsamplers do not change this, since they are trained to reproduce the backbone's encoding of the same field of view sampled more finely, which leaves the scale of analysis where the single global pass put it. We give a frozen DINOv3 a fovea. Conditioned on a point query, foveal re-encoding (FoRE) runs a lightweight modulator on the coarse map and the queried box of that same global view, predicting the features the backbone would produce by re-encoding the box at native resolution, in 2.8 milliseconds per query and almost six times more cheaply. It reads no pixel outside the view an image-guided upsampler receives, and what it learns is a remapping of that view's encoding rather than a reconstruction from pixels, so it differs from an upsampler in what it predicts rather than in what it reads. Distillation against the backbone's own re-encodings needs no labels or backbone updates. Across 25 datasets, more than fifteen readouts, and three domains, the predicted features are the strongest source apart from the teacher almost everywhere and lead every upsampling baseline by wide margins under in-context matching, with the largest gains where one wide-field encoding collapses most structure. On per-patch readouts they surpass even the re-encoding they were distilled from, an effect we trace to variance reduction and summarize in an averaging law for when it appears and reverses.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.