CompatLift: Anchored Feature Lifting for Remote Sensing Open-Vocabulary Segmentation
Abstract
Remote-Sensing Open-Vocabulary Semantic Segmentation (RS-OVSS) requires fine spatial detail to delineate small objects and narrow structures, while extracting high-resolution features with the backbone incurs substantial computation. Feature upsampling provides a more efficient alternative, but existing learned feature upsamplers are mainly developed for pretrained visual encoders and may produce representations that are incompatible with a frozen encoder–decoder segmentation pipeline. We study this problem in a SAM 3-based RS-OVSS pipeline and observe that directly replacing native features with learned upsamplers substantially degrades segmentation performance, whereas bicubic interpolation largely preserves the behavior of the frozen decoder. Motivated by this, we propose CompatLift, a lightweight feature upsampler framework that learns image-guided residual updates anchored by the bicubic reference. The upsampler module is trained without segmentation annotations using high-resolution features from the same frozen SAM 3 encoder, while the backbone and prediction heads remain frozen. We further introduce a simple score calibration step to compensate for the confidence shift caused by feature lifting. Extensive experiments on representative eight remote-sensing benchmarks show that CompatLift improves SAM 3 segmentation performance by lifted higher-resolution, decoder-compatible feature representations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.