CacheTok: One Binary State per Sampled Image Token, Read by Every Layer
Abstract
Block-autoregressive image generators such as VAR, Infinity, and PAR sample an image in a few blocks of tokens, yet every block attends to all earlier tokens, whose keys and values the transformer stores separately at each layer: Infinity-2B keeps 3.1 GiB of them for one image. Existing compressors for these models choose which tokens to keep, while later scales keep reading the oldest ones. We locate the redundancy along depth. Once a block is sampled, the per-layer keys and values of its tokens are views of one fixed content, and one linear subspace shared by all layers captures most of their energy. CacheTok writes each sampled token once, as sign bits per key–value head computed from its final hidden state and its token embedding, and every layer decodes its own keys and values from these bits through two linear readers; the losses of later blocks train the writer through a straight-through estimator. We prove that a shared state never needs more width than per-layer states and that exact training becomes sequential over blocks, a cost that grows with the number of blocks and stays small for next-scale generators with 10–13 blocks. On VAR-d16, PAR-XL4x, and Infinity-2B, CacheTok stores 512×, 1152×, and 1024× less visual history than matched full caches, changes FID by at most 0.03 and GenEval by −0.23 points, and doubles the batch capacity of Infinity-2B on one 80 GB GPU.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.