acceptodds
Under review as a conference paper at ICLR 2027

Same Recipe, Different Weights: Quantization and Decision Stability Across Backends

Abstract

Can one weight-only recipe be reapplied across tensor backends without changing its checkpoint, decisions, or cost? We evaluate symmetric INT8 and grouped INT4 on ten pretrained checkpoints spanning text, vision, audio, dual-encoder, and generative vision-language architectures in one multi-backend implementation on an A100. All 30 cross-backend FP32 state pairs match exactly; none of the 60 independently quantized pairs does. Exact PyTorch packed-state transplantation succeeds for eight ViT/CLIP target interventions and preserves all 800 held-out image decisions per intervention, yet a prespecified 99% calibration-agreement selector misses its holdout target on ViT and CLIP. Across three language-model families, all FP32 and INT8 16-token continuations agree across backends on eight fixed prompts, whereas minimum pairwise INT4 agreement is 5/8 on Qwen3 and Gemma3. Image-conditioned Qwen3-VL continuations agree across all backends for FP32 and INT8, but the worst INT4 pair agrees on only 15/20 images. Median stored-variable bytes fall to 0.254 and 0.184 of FP32 under INT8 and INT4, but portable direct-eager INT4 execution is slower across all three backends. These findings separate recipe reproducibility, exact state transfer, observed decision stability, and deployment cost without treating a changed tensor hash as a changed prediction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.