acceptodds
Under review as a conference paper at ICLR 2027

TASTE: Training-Aware Data Selection for Text-to-Image Generation

Abstract

The remarkable progress of text-to-image (T2I) generative models, such as Imagen, Stable Diffusion, and FLUX, has led to substantial improvements in visual quality. However, their performance is fundamentally limited by the quality of training data. Web-crawled and synthetic image datasets often contain low-quality or redundant samples which lead to degraded visual fidelity, unstable training, and inefficient computation. Hence, effective data selection is crucial for improving data efficiency. Existing approaches rely on costly manual curation or heuristic scoring based on single-dimensional features in text-to-image data filtering. Although meta-learning-based data selection has been explored for LLMs, its application to text-to-image model training remains underexplored. To this end, we propose **TASTE**, a meta-gradient-based framework to select a suitable subset from large-scale text-image data pairs. Our approach automatically estimates the influence of each sample by iteratively optimizing the model from a data-centric perspective. TASTE consists of two key stages: data rating and data pruning. We train a lightweight rater to estimate each sample's influence based on gradient information, enhanced with multi-granularity perception. We then use the Shift-Gsample strategy to select informative subsets. Experiments on both synthetic and web-crawled datasets demonstrate that Taste consistently improves downstream model performance. Training on the 50% subset selected by TASTE achieves performance comparable to full-data training while improving data efficiency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.