Dense Similarity as a Design Choice: Consistency and Utility in CLIP Readouts
Abstract
CLIP's native image–text function does not uniquely determine the internal similarities used by dense readers. We turn this freedom into Learned G, a dense adaptation method that learns task-directed metrics within a family of native-equivalent Q/K implementations. Reciprocal query/key transforms preserve native attention while changing the self-similarities used for dense aggregation. The transforms fold into existing projections without online adapter multiplications. On ViT-B/16, G improves matched SCLIP by 3.49 mIoU points over learned per-head scalars on VOC. Removing isotropic scale retains gains after scalar recalibration; spectrum-preserving rotations reduce performance. Retraining on ViT-L/14 yields gains over per-head scalars of 8.13 points on VOC and 4.89 on a fixed COCO-Object panel. With a trained output head held fixed, learning G upstream adds 1.97 points on VOC, showing that metric adaptation can complement output-side adaptation. Canonicalization gives qualifying readers consistent predictions across equivalent implementations; metric–bias transport retains a selected task across coordinate changes. Dense supervision can thus improve segmentation without changing CLIP's native function.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.