Structure over Pixels: Learning Variable-Length Visual Programs
Abstract
Discrete visual tokenizers map images to ordered sequences of tokens, providing a natural representation for structural scene descriptions. Most use a fixed sequence length, while adaptive methods often require post-hoc search or choose among a small set of rates that control the length. We propose STROP, a discrete tokenizer that learns both a visual program and its image-dependent active length. A length head is trained with a four-phase curriculum using local rate–distortion probes against frozen DINOv3 features, then predicts the active prefix in a single forward pass. At a matched rate of about nominal bits per crop, the adaptive model improves segmentation over a separately trained fixed-length baseline on four benchmarks (by – mIoU), and it also beats a fixed baseline that uses more bits. STROP programs also yield higher segmentation mIoU than FlexTok, One-D-Piece, and ALIT at similar or higher rates, under the same readout architecture and training protocol. STROP therefore learns useful per-image sequence lengths without post-hoc search or a predefined set of compression rates.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.