acceptodds
Under review as a conference paper at ICLR 2027

Dense Similarity as a Design Choice: Consistency and Utility in CLIP Readouts

Abstract

CLIP's native image–text function does not uniquely determine the internal similarities used by dense readers. We turn this freedom into Learned G, a dense adaptation method that learns task-directed metrics within a family of native-equivalent Q/K implementations. Reciprocal query/key transforms preserve native attention while changing the self-similarities used for dense aggregation. The transforms fold into existing projections without online adapter multiplications. On ViT-B/16, G improves matched SCLIP by 3.49 mIoU points over learned per-head scalars on VOC. Removing isotropic scale retains gains after scalar recalibration; spectrum-preserving rotations reduce performance. Retraining on ViT-L/14 yields gains over per-head scalars of 8.13 points on VOC and 4.89 on a fixed COCO-Object panel. With a trained output head held fixed, learning G upstream adds 1.97 points on VOC, showing that metric adaptation can complement output-side adaptation. Canonicalization gives qualifying readers consistent predictions across equivalent implementations; metric–bias transport retains a selected task across coordinate changes. Dense supervision can thus improve segmentation without changing CLIP's native function.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.