acceptodds
Under review as a conference paper at ICLR 2027

When Steering Works: Predicting Debiasing Benefit from Denoiser Responses

Abstract

Semantic steering can reduce stereotypes in text-to-image generation, but it can also change content that the prompt asks for. Can we predict which prompts will benefit before decoding their images? We introduce Directional Gain (), which compares two denoiser responses at the same latent state: the change caused by steering and the change caused by replacing a stereotype prompt with its counterfactual. We measure agreement between these directions relative to random edits. Across SD3.5, FLUX, and Qwen-Image, with 1,295 test prompts per model, predicts improvement better than response size. At the top 20%, selection by gives larger stereotype reductions and usually smaller prompt-matching losses. Its advantage in combined-score improvement over response size is 4.7–20.7 points per selected prompt under Qwen3-VL; GPT-5.6 also favors in all three models. Averaged over six single-probe choices, one probe retains 89–93% of the mean improvement obtained by six-probe selection. It uses fewer denoiser evaluations than generating a clean–steered image pair and requires no image decoding or judging. With the steering settings fixed, the ranking also selects prompts that benefit on new generation seeds.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.