Can Language Represent Semantic Variation? A Semantic Basis for VLM Robustness Audits
Abstract
Vision-Language Models (VLMs) underpin a broad range of real-world AI applications. Yet understanding their robustness under semantic variation remains difficult, since such variations are hard to formalize directly in image space. We therefore ask whether language can provide a reliable representation of semantic variation for robustness auditing. We construct semantic directions from semantic-contrast text pairs and aggregate them across textual contexts for consistency. Cross-modal correction further calibrates these directions to the visual embedding space. Controlled visual validation shows that language-derived semantic representations align well with corresponding visual semantic variations. Building on this semantic representation, we establish a semantic robustness auditing framework for VLMs that quantifies prediction stability along semantic variation with closed-form boundary margins. The framework supports auditing across data, prompt, and model contexts, while enabling fine-grained diagnosis of where robustness is preserved or degraded. Our audits reveal systematic variation in semantic robustness across contexts alongside stable semantic patterns. Domain-specific applications further show that the audit localizes fragile patterns aligned with task priors and guides model selection under target-domain robustness requirements. Overall, our framework provides a principled basis for evaluating and diagnosing VLM robustness under semantic variation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.