TrajQuant: Trajectory-Aware Quantization of Interactive Video World Models
Abstract
Interactive video world models generate future visual states conditioned on actions and reuse those states as history for subsequent interaction. Their repeated denoising and long autoregressive rollouts make low-bit inference attractive, but conventional local reconstruction does not explicitly account for how quantization errors propagate along action-conditioned trajectories. We study this mismatch as a **control-to-trajectory error** problem: a local perturbation introduced during one chunk alters the inputs encountered by later chunks, while layer-wise tensor reconstruction may fail to capture errors in the complete state transition. We propose **TrajQuant**, a trajectory-aware post-training quantization framework for action-conditioned autoregressive video world models. Its **MLGR** module applies structured rotations whose angles are refined against complete quantized Transformer-block outputs, improving reconstruction at the transition level. Its **FGCC** module (Fisher-Guided Chunk Calibration) estimates a squared-gradient sensitivity proxy for each chunk–layer pair and uses it to weight paired full-precision and upstream-quantized reconstruction, accounting for errors propagated from preceding chunks. MLGR is fixed before FGCC performs final weight rounding in the transformed coordinates. We evaluate Matrix-Game-2 and ABot-World-0 at W6A6, W4A8, and W4A6 with WBench, LPIPS-FP16, and FVD-FP16. On Matrix-Game-2 at W4A6, TrajQuant achieves LPIPS-FP16/FVD-FP16 of 0.1294/32.71 versus SVDQuant’s 0.1709/45.26. On ABot-World-0 at W4A6, it also achieves the lowest LPIPS-FP16 and FVD-FP16 among the listed quantized methods, while the leaders at other precisions differ by metric.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.