Do Models Share Safety Representations? Cross-Model Steering for Safe Visual Generation
Abstract
Recent progress in generative modeling has made safety control a central challenge, yet existing approaches remain largely model-specific, requiring retraining or tailored interventions for each new architecture. In this work, we introduce cross-model safety steering, a framework that extracts a safety direction from paired safe and unsafe prompts in a source LLM and maps it to the text-conditioning space of a target generator through an alignment learned exclusively from benign data. Once transferred, the direction is applied at inference time while keeping the generative backbone frozen and without requiring target-side unsafe examples for alignment. The framework also supports category-specific directions, enabling selective control over different types of unsafe content. Across diverse source-target pairs for text-to-image and text-to-video generation, transferred directions consistently reduce unsafe outputs, with the alignment strategy and steering strength determining the balance between safety and generation quality. We further show that steering strength can be selected using benign data alone, reducing unsafe outputs in all 45 evaluated configurations. Together, these results show that safety interventions can transfer across independently trained models, offering a modular alternative to model-specific steering. Source code will be made publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.