acceptodds
Under review as a conference paper at ICLR 2027

Beyond Refusal: Building and Evaluating Diffusion Post-hoc Safety Intervention Models

Abstract

Existing safety mechanisms primarily rely on refusal or rewriting to block harmful outputs. However, in real-world scenarios where safe and harmful content are intertwined, discarding or rewriting the entire response is not only inefficient but also results in the loss of valuable information. To address this, we shift our focus to post-hoc safety interventions. The effectiveness of a post-hoc intervention model should be evaluated on its ability to accurately identify risks, precisely edit harmful spans, and preserve benign context, all while maintaining high inference efficiency. To excel across these dimensions, we introduce SafeLLaDA, a diffusion-based safety model. By leveraging the unique iterative refinement properties of diffusion models, SafeLLaDA precisely locates and modifies harmful spans post-generation while leaving the benign context unchanged. Experimental results demonstrate that SafeLLaDA significantly mitigates harm and achieves superior content fidelity compared to baseline guardrails and other traditional safety measures, establishing diffusion-based post-hoc intervention systems as a highly viable alternative to traditional refusal or rewrite strategies.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.