acceptodds
Under review as a conference paper at ICLR 2027

Structure over Pixels: Learning Variable-Length Visual Programs

Abstract

Discrete visual tokenizers map images to ordered sequences of tokens, providing a natural representation for structural scene descriptions. Most use a fixed sequence length, while adaptive methods often require post-hoc search or choose among a small set of rates that control the length. We propose STROP, a discrete tokenizer that learns both a visual program and its image-dependent active length. A length head is trained with a four-phase curriculum using local rate–distortion probes against frozen DINOv3 features, then predicts the active prefix in a single forward pass. At a matched rate of about nominal bits per crop, the adaptive model improves segmentation over a separately trained fixed-length baseline on four benchmarks (by – mIoU), and it also beats a fixed baseline that uses more bits. STROP programs also yield higher segmentation mIoU than FlexTok, One-D-Piece, and ALIT at similar or higher rates, under the same readout architecture and training protocol. STROP therefore learns useful per-image sequence lengths without post-hoc search or a predefined set of compression rates.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.