COLOR ANNOTATIONS IN TEXT-TO-IMAGE PRETRAINING: AN EMPIRICAL STUDY
Abstract
Numerical color annotations provide a route to precise color control in text-to-image pretraining, but how their design affects generation with and without explicit color inputs remains unclear. We investigate this question by training diffusion transformers from scratch on a shared BLIP3-o subset, keeping the architecture and training hyperparameters fixed. We compare region-level palettes extracted from image pixels with object colors predicted jointly with color-enriched descriptions. With benchmark colors explicitly supplied, our pixel-annotated model reaches 86.97% accuracy on the evaluated GenColorBench numerical-color task. Its scores exceed published NumColor results in the three color categories compared, under the respective evaluation settings. Within our palette-trained configurations, pixel-extracted annotations give the strongest supplied-color fidelity, including prompts requiring multiple colors on multiple objects. This ranking reverses when palettes are withheld from matched prompts while descriptions, target colors, and scoring remain fixed. The pixel-annotated model also gains more in image quality from receiving palettes, whereas predicted annotations support stronger texture binding and compositional generation. These results demonstrate the effectiveness of color-annotated pretraining for precise numerical control. They also show that, in this comparison, the strongest control of supplied colors accompanies the greatest palette dependence and does not translate into the strongest color generation from descriptions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.