acceptodds
Under review as a conference paper at ICLR 2027

Understanding and Enhancing Conditioning Robustness of Text to Image Diffusion Models

Abstract

Recent advances in diffusion models have significantly improved the quality of text-to-image generation. Yet visual fidelity alone is not sufficient: reliable conditioning is essential for predictable control, and its limitations continue to constrain the broader success of these models. In particular, the sensitivity of generated images to non-semantic perturbations of the input prompt is a widely recognized problem. In this work, we develop a simple perturbative model of the mechanism underlying this instability and validate it through targeted experiments. We further introduce dedicated metrics and a benchmark (CoRBen) for systematically evaluating prompt robustness. The resulting theory and empirical findings motivate the sensitivity measure and a family of training-free plug-and-play methods that improve generation stability at inference overheads of 100%, 17%, and 3%. Our code is available at https://anonymous.4open.science/r/conditioning-robustness/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.