Global FP4 Quantization via Low-Dimensional Convex Optimization
Abstract
Native FP4 quantization encodes small weight groups with shared scales, but correlated inputs couple their reconstruction errors. Independent groupwise selection neglects these interactions, while exhaustive joint search grows exponentially with the number of groups. For reconstruction metrics with full within-group curvature and a shared rank- component, we formulate a convex relaxation with an -dimensional dual. We prove that any feasible candidate mixture can be reduced to at most mixed groups without changing its aggregate residual or increasing the relaxed objective. LDC-FP4 uses this construction to coordinate codes and scales and recover native encodings with a residual-variance bound. Across Llama and Qwen models, LDC-FP4 improves language modeling and downstream performance over the evaluated quantization baselines. Matched-candidate comparisons show that global coordination improves reconstruction beyond independent selection and outperforms a wide-beam search baseline at significant lower calibration cost. All optimization is confined to calibration, leaving the native FP4 representation and inference operators unchanged.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.