acceptodds
Under review as a conference paper at ICLR 2027

DiveUp: Learning Feature Upsampling from Diverse Vision Foundation Models

Abstract

Recently, feature upsampling has gained increasing attention owing to its effectiveness in adapting vision foundation models (VFMs) for pixel-level understanding tasks. While backbone-agnostic upsamplers offer lower training overhead and broader applicability, existing methods suffer from two fundamental limitations. First, relying solely on single-model intra-reconstruction forces the upsampler to overfit to the source model's inherent spatial misalignments and high-norm artifacts. Second, it remains unclear how to compose the training backbones for universal zero-shot generalization. To address these limitations, we propose DiveUp, a unified upsampling framework that breaks single-model dependency via two complementary mechanism: training on a diverse mix of VFMs, and supervision the upsampler with a cross-model relational guidance derived from a single spatially-aligned teacher. Concretely, Diveup formulates a local center-of-mass (COM) field as a universal relational representation that transfers the teacher's geometric structure to correct the student's spatial interpolation, without altering the backbone's native semantics. Furthermore, through systematic ablations across 14 diverse VFMs, we uncover two distinct principles governing zero-shot generaliztion: (1) at a fixed backbone budget, the representation diversity of the training mix, rather than the identity of any bacbbone, governs out-of-distribution geralizablitiy; and (2) explicit spatial alignment guidance is a separate regularization that is decisive for structurally noisy backbones. Extensive experiments on dense prediction demonstrate that DiveUp achieves state-of-the-art zero-shot performance across unseen architectures.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.