acceptodds
Under review as a conference paper at ICLR 2027

ThunderWorld: Efficient World Model Generation with Selective FP4 Refinement via Native GEMMs

Abstract

We study whether world models can recover high-quality generation through dense four-bit floating-point (FP4) matrix computation without sparse matrix instructions or token/channel reordering. ThunderWorld is a training-free framework that corrects four-bit weight-and-activation (W4A4) error using additional dense FP4 products formed from quantization residuals. Block-aligned supports jointly govern residual storage, online preparation, and dense execution under a compute budget, preserving every base attention block. Controlled attention diagnostics support structure-informed selection over random and window policies at matched repair computation. On Cosmos3 Edge and Nano across RBench and PAI-Bench-G, a configuration capped at twice base W4A4 matrix work keeps mean task-score gaps within 0.08 RBench points and 0.37 PAI-Bench-G percentage points of BF16. It reduces LPIPS, a perceptual distance to same-input, same-seed BF16 generations, by 3.3–7.8% versus ARCQuant (2.0), with 1.17–1.88× its throughput on RTX PRO 6000 and Jetson AGX Thor. These results provide a numerical and systems basis for world-model inference on devices with native FP4 Tensor Cores.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.