CoordUp : Upsampling Any Feature Beyond the Megapixel Regime
Abstract
Vision foundation models produce semantically rich but spatially coarse features, limiting fine-detail recovery in dense prediction tasks. Existing feature upsamplers tackle this problem but are often tied to specific backbones or require materializing dense feature grids, making high-resolution inference costly. We introduce CoordUp, a lightweight, feature-agnostic upsampler that formulates feature upsampling as a coordinate-conditioned operator. Across five frozen backbones and three high-resolution datasets for semantic segmentation and depth estimation, it consistently improves boundary-sensitive metrics over previous encoder-agnostic upsamplers. It also transfers zero-shot to the final activations of DA3 and SAM3, improving boundary quality without retraining. Finally, its coordinate-based formulation allows CoordUp to scale gracefully with output resolution, enabling, for example, an 80 Mpx output using only 8.6 GB of GPU memory.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.