acceptodds
Under review as a conference paper at ICLR 2027

Channel-group Hierarchical Allocation for Mixed Precision

Abstract

Low-precision quantization can substantially reduce the memory and computation costs of Transformer inference, but aggressive 4-bit floating point (FP4) quantization often degrades model quality due to highly non-uniform quantization sensitivity across the model. Mixed-precision quantization addresses this challenge by assigning higher precision to sensitive regions while keeping less sensitive regions in lower precision. Existing fine-grained mixed-precision methods often select high-precision regions directly from local or global sensitivity scores, which can over-concentrate the high-precision budget in a subset of the model while leaving other sensitive regions insufficiently protected. We introduce CHAMP (Channel-group Hierarchical Allocation for Mixed Precision), a post-training framework combining FP4 and 8-bit floating point (FP8) that hierarchically distributes a global high-precision budget across model components and Transformer layers and constructs the resulting fine-grained channel-group precision mask using sensitivity-aware selection. CHAMP combines sensitivity- and size-aware module allocation, layer-specific adjustment, and local–global channel-group selection. Across language and video generation tasks, CHAMP provides the largest improvements when the FP8 budget is limited, outperforming prior fine-grained sensitivity-based mixed-precision allocation in this regime. CHAMP narrows the gap to full FP8 in our primary language-model settings, improves video frame-level fidelity, maintains comparable generation quality, and achieves higher theoretical throughput.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.