Semantically Steerable Image Tokenization
Abstract
Adaptive image tokenizers compress an image into a variable number of tokens matched to its complexity, and recent work makes this allocation controllable: conditioned on a target reconstruction error, the tokenizer itself decides how many tokens an image needs. However, the fidelity measure guiding this allocation shapes which aspects of the image are preserved under compression. The pixel-wise error used in prior work is blind to the spatial structure that carries an image's meaning, and under heavy compression such tokenizers are prone to abruptly discarding tokens that encode essential content. We introduce GIST (**G**uided **I**mage **S**emantic **T**okenizer), an adaptive tokenizer that can be steered toward a requested *semantic* fidelity. Its token allocation is jointly controlled by pixel-level and semantic target errors, with semantic error measured in the feature space of a frozen DINOv3 encoder. This joint control enables *semantic protection*: with the semantic target fixed at zero, the reconstruction target can be relaxed to discard pixel detail while preserving what the image depicts. In a controlled comparison on ImageNet-100, GIST outperforms its pixel-controlled counterpart KARL at every token count, with up to % lower LPIPS and % lower DreamSim. On average, lowering its semantic target reduces LPIPS faster per token than lowering the reconstruction target, and up to in the low-token regime. Semantic fidelity is thus an effective control signal for adaptive tokenization: it lets a tokenizer see, and so preserve, what an image depicts.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.