AlignLedger-CLIP: Group-Wise Auditing and Guarded Adaptation for Frozen Vision-Language Embeddings
Abstract
Aggregate retrieval scores can conceal regressions within semantic groups when adapting frozen vision–language embeddings. We introduce ALIGNLEDGER-CLIP, a framework that connects group-wise auditing to adaptation and update selection. A ledger records group geometry, directional retrieval, and representation changes; its statistics weight gap-closing and preservation objectives for lightweight cached- feature adapters. A held-out guard selects an interpolation scale subject to recall and group-degradation constraints, with the frozen representation as a fallback. The same ledger then measures the deployed update on evaluation records. We study 12 evaluation slices from 10 source datasets with CLIP and SigLIP, separating transductive audits from disjoint fit/guard/evaluation experiments. In the latter, guarded ALIGNLEDGER-CLIP achieves 38.27 average Recall@1 with 20.01% degraded groups, while identically guarded CLIP-Refine achieves 38.53 with 21.65%. These retrieval–degradation profiles expose the local cost of adaptation alongside its aggregate benefit.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.