acceptodds
Under review as a conference paper at ICLR 2027

Safe Vision-Language Models via Sparse Autoencoder Feature-Guided Embedding Steering

Abstract

Large vision language models such as CLIP provide transferable representations for retrieval, classification, generation, and visual question answering. However, pretraining on web data can also encode unsafe concepts that propagate to downstream systems. Existing safety alignment methods mainly follow two strategies. Fine tuning with constructed safe and unsafe pairs requires substantial training and can disturb broadly useful knowledge. Methods that directly manipulate pretrained weights without further training avoid this cost, but often provide limited safety control and may still impair general capabilities. We attribute these limitations in part to unsafe and general semantics being entangled in dense representations. Inspired by sparse autoencoders, which decompose dense neural representations into interpretable features, we propose Sparse Autoencoder Feature Guided Embedding Steering (SAFES-CLIP). Rather than updating the pretrained backbone, our method trains a sparse autoencoder with multiple Top-K budgets on frozen embeddings and obtains a shared sparse concept space. It identifies features associated with unsafe content using safe and unsafe calibration samples, then removes only their contribution to each input during inference. We further extend this mechanism to token representations for integration with generation and visual question answering systems. Experiments across retrieval between images and text, image classification, image generation, and visual question answering show consistent safety improvements while preserving broad general capabilities. Our core code has been provided in the supplementary material.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.