acceptodds
Under review as a conference paper at ICLR 2027

See More from Less: Sample-Efficient Generative Feedback for Fine-Grained Visual Understanding

Abstract

Generative feedback can improve the fine-grained visual understanding of Contrastive Language-Image Pre-training (CLIP), but existing methods require substantial data and computation. We identify two sources of sample inefficiency: first, projected representations can collapse toward a shared constant, weakening gradient feedback to CLIP; second, hierarchical information is underutilized, as supervision from only one or a few levels limits feedback per sample, while directly combining multiple levels can introduce conflicting gradients. We propose Sample-efficient Generative Feedback (SGF) to extract more useful supervision from each training sample. Margin regularization encourages matched conditions to yield lower reconstruction losses than mismatched ones by a positive margin, preserving image-specific information. Hierarchical boosting progressively integrates coarse-to-fine supervision while retaining earlier objectives, giving them time to stabilize before finer supervision is introduced, thereby reducing cross-level interference. Experiments across six CLIP backbones show that SGF achieves state-of-the-art fine-grained perception using only 84K training samples, at least 82% fewer than prior methods, with a - training speedup over the fastest competitor, demonstrating its improved sample efficiency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.