acceptodds
Under review as a conference paper at ICLR 2027

Can LLMs Understand and Generate Images Pixel by Pixel?

Abstract

Large language models (LLMs) can generate images through text formats such as SVG and CAD, but these representations use high-level primitives rather than explicit raster values. We ask whether LLMs can instead generate pixel images directly with their existing vocabulary, without a learned image tokenizer or decoder. To test this, we define a palette-indexed representation: a model outputs an RGB palette and a grid assigning one palette color to each pixel. Using it, PixelBench evaluates text-to-pixel (T2P) and image-to-pixel (I2P) generation across resolutions, alongside a complementary pixel-understanding task. Models show substantial pixel-generation ability in both tasks. Cross-resolution analysis reveals model-specific scaling: most models improve and then decline with resolution, whereas the strongest produces high-quality, recognizable images and remains robust. Pixel-understanding performance follows a similar model ranking and correlates strongly with generation performance. This connection motivates a second question: can training models to reconstruct pixels improve general visual understanding? To investigate, we introduce Pixel Reconstruction Tuning (PRT), a self-supervised objective that trains VLMs to reconstruct images as pixel grids. PRT improves average performance across 13 visual-centric benchmarks}, indicating that explicit pixel reconstruction can provide useful auxiliary supervision for visual understanding.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.