Auditing VLM Repair: Variance Geometry and Guarded Candidate Selection
Abstract
Vision-language models (VLMs) are often repaired at deployment by changing the geometry of frozen image and text embeddings. Whitening, prompt adaptation, and nuisance removal can improve average accuracy, yet the same operation can amplify a spurious factor and damage the least represented groups. We study this problem as an audit and selection task: a repair method proposes candidates, while a deployment layer decides which candidate is safe to apply. We introduce a group-free guarded selector built from task-conditioned variance diagnostics, validation-average preservation, and prompt stability. The diagnostics separate semantic alignment from nuisance alignment in the observed covariance spectrum and expose when covariance corrections should be used, ignored, or filtered. Across 156 scenarios spanning Waterbirds, CelebA, synthetic CIFAR shifts, PACS, and two VLM backbones, naive validation-average selection yields a worst-group/domain delta of -0.006, whereas the guarded policy yields +0.037 with a 95% bootstrap interval of [+0.003, +0.078] and reduces harmful selections from 30 to 9. The combined risk signal identifies harmful observed-covariance candidates with AUROC 0.956. Adding TIE-style candidates to the pool raises robust delta to +0.073 and gives a paired gain of +0.041 over our candidates alone. The results establish deployment-time auditing as a complementary layer for VLM repair, while identifying a clear boundary where strong nuisance shift remains difficult.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.