Making Language the Low-Rate Code of Vision: Successive Refinement from Words to Pixels
Abstract
An ordered visual code lets a short prefix stand in for an image, and the training loss decides what that prefix holds. Under pixel reconstruction the first bits go to the highest-variance directions: we construct a source for which a code that is rate-optimal for pixels at every length carries no caption information in its entire first half. A caption is therefore not the low-rate code of its image unless the code is built to make it one. We cast captioning and text-to-image generation as successive refinement under a distortion measure that changes along the code, from caption log-loss in the early stages to pixel distortion in the last. Putting words first then costs nothing among the caption stages exactly when the captions form a summarization chain, and costs the pixels at most the caption uncertainty left in a pixel-optimal reconstruction, which is large for short prefixes and small at full length. We build RUNG, an ordered code whose prefixes are supervised by a summarization ladder of captions and whose full length reconstructs the image; generation extends the code past the prompt and captioning truncates it. With data and capacity held fixed, ordering by language instead of reconstruction raises prefix–caption information 2.9× and GenEval from 0.76 to 0.83, at a cost of 2.8 dB PSNR at a quarter of the code and 0.1 dB at full length. RUNG-XL (3.4B) reaches 0.88 GenEval and 87.6 DPG-Bench in 32 network evaluations, and 0.91 GenEval when errors are redrawn at the rung where they appear; its fine captions entail 94.1% of the attribute tuples in its coarse ones, and it edits images without inversion.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.