acceptodds
Under review as a conference paper at ICLR 2027

WAMPTQ: Beyond Weight Quantization via Task-Utility-Guided Precision Allocation for Efficient World Model Deployment

Abstract

World action models (WAMs) enable action-conditioned prediction and planning, but efficient deployment faces two challenges. Low-bit sensitivity varies across tasks, while reconstruction proxies do not directly measure native task utility. Prediction and context processing also create large transient tensors, so smaller packed artifacts need not reduce peak memory or latency. We introduce WAM post-training quantization (WAMPTQ), which couples task-utility-guided precision selection with liveness-aware execution. WAMPTQ constructs legal 4- and 8-bit weight maps under a matched packed-size budget and selects the final map using native task utility. An optional large language model (LLM) Agent narrows the legal proposal region before scoring. WAMPTQ then freezes the artifact and applies quantized epilogue fusion (QEF) and virtual-context consumer fusion (VCF) to remove avoidable rescaling and context-assembly intermediates. We separate within-pool selection error from candidate-coverage loss and derive an exact live-set criterion for when local savings reduce the global memory peak. Across 15 tasks from four WAM families, the largest observed quality drop relative to FP32 is 2.22 percentage points, with median end-to-end speedup and peak-memory reduction of and , respectively. The code will be released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.