acceptodds
Under review as a conference paper at ICLR 2027

From Knowledge Gaps to Better Generation: Auditing and Completing World Knowledge in Text-to-Image Training Data

Abstract

Text-to-image (T2I) models achieve high visual quality but remain inconsistent in depicting concepts grounded in world knowledge, particularly those sparsely represented in training data. Because datasets are typically organized by source, concept coverage gaps are difficult to identify and address. We present CGTR, a framework that maps large scale image–text data to a hierarchical world knowledge graph for coverage analysis and targeted training-data refinement. Auditing 370M annotation records from publicly available T2I data reveals missing and underrepresented concepts across semantic levels. These statistics guide the construction of KnowGapBench, which covers 3,286 underrepresented concepts and evaluates concept identity separately from compositional compliance. CGTR independently audits each target training dataset to identify supplementation targets. World knowledge and diverse scene graph templates guide caption construction, while curated reference images guide appearance during image synthesis. Caption consistency and image–text alignment checks filter the resulting training pairs. The hierarchy organizes original and supplementary data into semantic groups, and concept-aware sampling regulates their exposure during training. We apply CGTR to Fine-T2I and fine-tune UniWorld-V1 with the refined data under its original generative objective. The overall Knowledge score on KnowGapBench increases from 15.3 to 22.5, while performance on familiar concepts outside the supplementation target set is preserved.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.