Can LLMs Understand and Generate Images Pixel by Pixel?
Abstract
Large language models (LLMs) can generate images through text formats such as SVG and CAD, but these representations use high-level primitives rather than explicit raster values. We ask whether LLMs can instead generate pixel images directly with their existing vocabulary, without a learned image tokenizer or decoder. To test this, we define a palette-indexed representation: a model outputs an RGB palette and a grid assigning one palette color to each pixel. Using it, PixelBench evaluates text-to-pixel (T2P) and image-to-pixel (I2P) generation across resolutions, alongside a complementary pixel-understanding task. Models show substantial pixel-generation ability in both tasks. Cross-resolution analysis reveals model-specific scaling: most models improve and then decline with resolution, whereas the strongest produces high-quality, recognizable images and remains robust. Pixel-understanding performance follows a similar model ranking and correlates strongly with generation performance. This connection motivates a second question: can training models to reconstruct pixels improve general visual understanding? To investigate, we introduce Pixel Reconstruction Tuning (PRT), a self-supervised objective that trains VLMs to reconstruct images as pixel grids. PRT improves average performance across 13 visual-centric benchmarks}, indicating that explicit pixel reconstruction can provide useful auxiliary supervision for visual understanding.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.