acceptodds
Under review as a conference paper at ICLR 2027

COOPER: CLASS-OUTPUT OPTIMIZATION VIA PER-EXAMPLE EMBEDDING REFINEMENT

Abstract

Methods that learn class embeddings for frozen vision–language models typically optimize them jointly across labeled examples. We compare this approach with local optimization followed by aggregation: fitting one embedding per example, then combining the fits within each class. COOPER builds a reusable class bank from these local refinements; its classification counterpart, COOPER-CLS, fits the bank jointly. Both freeze model weights, select banks using training data, and deploy fixed banks without test-time optimization. Our theory explains why the two approaches can produce different embeddings and when weighting can reduce bias from ineffective refinements. We certify that no convex combination of the sampled text-prompt embeddings reproduces the tested COOPER corrections. We evaluate COOPER on eleven segmentation datasets with SAM3 and COOPER-CLS across nine classification backbones. COOPER yields positive gains over SAM3's original embeddings on all eleven datasets, averaging +[mean gain] mIoU points, and its banks also transfer to other datasets and video. Although controlled comparisons give dataset-dependent local–joint rankings, COOPER outperforms the tested fixed-recipe joint-fitting baselines on average, including joint fitting initialized with COOPER embeddings. For cosine classification with CLIP, PE, and SigLIP, however, our controls favor joint fitting over local aggregation. Our work shows that the choice between local aggregation and joint fitting shapes how well learned class embeddings generalize.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.