acceptodds
Under review as a conference paper at ICLR 2027

Aggregating Safety Guidance for Co-occurring Concepts via Quadratic Programming

Abstract

Despite advances in safety guidance for text-to-image generation, suppressing multiple co-occurring undesired concepts remains challenging. Prior work has shown that naively aggregating category-specific safety guidances can cause directional attenuation, motivating selecting the most aligned guidance with the current generative state. Although such heuristic can mitigate interference among concurrent safety signals, it remains unclear whether it can reliably suppress all targets when multiple undesired concepts co-occur within a single image. To investigate this question, we introduce CHC-Bench, a benchmark for evaluating the simultaneous suppression of restricted concepts of different kinds. We construct a prompt-generation and verification pipeline that retains prompts whose generated images contain all specified targets, enabling evaluation of their joint suppression. We further propose QPS, which formulates multi-guidance aggregation as a quadratic program. QPS finds the minimum-norm correction to the base generative update while jointly preserving each active guidance source's standalone directional effect. Experiments on CHC-Bench demonstrate the importance of jointly considering multiple guidance signals for suppressing co-occurring targets. Evaluations on conventional safety benchmarks further demonstrate a balanced safety–utility tradeoff, indicating that joint aggregation remains useful in single target settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.