acceptodds
Under review as a conference paper at ICLR 2027

Unspoken, Controllable, Unavoidable: Implicit Co-occurring Objects is A Fundamental Property in Text-to-Image Generation

Abstract

Text-to-Image (T2I) models often generate objects that are not in the prompt: e.g., a cup on a table when the prompt is simply “a photo of a cup”. We call those the Implicit Co-occurring Objects (ICO). In this work, we present the first systematic study of ICO across seven architecturally distinct T2I models spanning the years 2022-2026 with a strategic plan for the understanding and the control of them. We find that 34-50% of the generated images featuring common objects (56,000 across 80 categories), regardless of the model, contain at least one unprompted object, with strikingly consistent co-occurring patterns shared across the models. To understand where they come from and how models generate them, we investigate four hypotheses about the training data, text encoders, latent space trajectory planning, and sampling dynamics. Training data co-occurrence statistics are found to be the dominant factor, and the text encoder has a secondary effect. And the sampling parameters modulate the magnitude. Guided by the insights, we explore a group of controls of ICO and discover a simple yet effective two-component control method combining isolation phrasing with embedding steering that reduces the leakage rate to 0.19% with a 94.2% fidelity. For a theoretical understanding, we show that having ICO is a fundamental property of the training data, labels, and the objective of T2I models: with co-occurrence statistics existing in the natural image distribution, there is a trade-off between the model-vs-data distribution divergence and the rate of ICO, meaning we cannot freely suppress ICOs without a loss of generation fidelity. Finally, we show that ICO carries a downstream cost: object detectors trained on T2I-generated images exhibit up to an 18.8% gap in recall compared to ICO-free data. We are the first to establish a comprehensive understanding of ICO in T2I models. Our findings provide new insight into this intrinsic property of diffusion-based image synthesis.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.