CSOA: Concept-Selective Adaptation with Semantic Control for Vision-Language Unlearning
Abstract
Concept unlearning in vision-language models (VLMs) is increasingly critical for selective removal of unwanted semantic information. However, practice demonstrates that while existing methods report success on traditional metrics, target recognition may remain recoverable through unconventional queries or visual domains. We identify three critical failure modes: visual bias arises when unlearning on narrow forget sets leaves recognition intact in unseen visual domains; association leakage occurs when suppressing specific text templates leaves alternative expressions as retrieval backdoors; and semantic collapse results when unconstrained adaptation disrupts meaningful post-unlearning predictions and neighboring concepts. To address these challenges, we introduce CSOA, a concept-selective adaptation framework with semantic control. This approach combines a quadrant anchor system specifying the linguistic neighborhood with hierarchical control that guides edited representations toward user-specified semantic ancestors, while preserving non-target utility. Knowledge-matched controls and factorial interventions assess semantic redirection and the retention-focused variant CSOA-R. Tests on three CLIP backbones and 40 additional concepts examine breadth; paired generative comparisons distinguish recognition suppression from answer disclosure.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.