acceptodds
Under review as a conference paper at ICLR 2027

Position-Sensitive Concept Acquisition in Text-to-Image Models

Abstract

Pretrained text-to-image models possess rich visual priors, yet reliable control over quantitative attributes can remain difficult to acquire through supervised fine-tuning. We study how the position of a target condition affects this acquisition process. Across models, concepts, and evaluation settings, paired fine-tuning with identical image supervision and matched training budgets yields markedly different control capabilities when the same target field is placed at the beginning or end of the prompt. The effect depends on both the model and the concept. Using angle control with Qwen-Image as a representative case, we find evidence linking the acquisition gap to information propagation in the frozen causal text encoder and the use of target information during training. A simple Prefix fine-tuning scheme repeats the target condition at the beginning and improves control over standard end-position training when the prefix is retained at inference. We further introduce a pre-adaptation probe of relative preference for correct over counterfactual conditions. Cross-model comparisons and controlled pretraining experiments illustrate three characteristic acquisition patterns under finite budgets: weak control at both positions, a pronounced positional gap, and stronger control at both positions. These patterns are associated with initial support and learning progress, rather than fixed support thresholds. Acquired control may remain format-dependent even as positional gaps narrow.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.