Who Needs More Bits? Task-Aware Quantization for Multi-Source Collaborative Inference
Abstract
Collaborative AI models increasingly integrate multiple views, modalities, or agents to solve a shared task, yet different information sources can contribute unequally to the final prediction. This raises a fundamental question for model quantization: how much precision should each source receive, and what determines its precision requirement? We study this problem from an information-theoretic perspective. We first establish a local-to-global perturbation characterization that connects source-wise parameter quantization errors to collaborative output distortion. Building on this result, we derive lower and upper rate-distortion bounds for task-aware multi-source model quantization by jointly accounting for source-specific parameter statistics and task contributions. We further characterize heterogeneous source contributions from a partial information decomposition perspective. Based on these results, we derive an explicit precision-allocation law and develop a lightweight precision-allocation rule that maps the theoretical solution to hardware-supported bit-widths. Experiments across multiple collaborative models, quantization methods, and practical resource constraints show that the derived bounds capture practical performance trends and effectively guide source-wise precision allocation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.