acceptodds
Under review as a conference paper at ICLR 2027

Tokenize the Physics, Not the Pixels: TessTok and Fold-Flow Quantization for Material Microstructures

Abstract

Visual tokenizers compress images into the tokens that latent generative models learn from. Existing tokenizers are built for natural images: a fixed latent grid, a codebook or straight-through estimator, and a perceptual objective. A physical field is instead a set of objects that a simulation reads directly. *Can a discrete tokenizer compress a material microstructure field by two orders of magnitude and return a field a crystal-plasticity solver can run?* We present **TessTok**, which encodes a polycrystal as sparse quantized Laguerre grain seeds and regrows it deterministically. Its **Fold-Flow Quantizer (FFQ)** needs no codebook or straight-through estimator and learns where to place its levels. At kbit per field reconstruction, TessTok reaches dB and rFID and matches of the grains, against dB, and for ViT-VQGAN at kbit. At matched bits, FFQ beats every codebook and binary quantizer on all pixel and orientation metrics with no codebook to collapse. TessTok also keeps of grain area at kbit, below the rate of MAGVIT-v2. On ImageNet, with its grain-specific parts switched off, the same code gives the highest PSNR among the latest tokenizers, though not the best perceptual scores and compression rate. Tokenizing physical-object fields instead of pixels, lets our model work at many bit budgets, and lets a reconstruction be governed by the physical simulation needs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.