Autonomous Model Quantization with Explorer–Reviewer Agents
Abstract
Model quantization is critical for efficient deployment of large language models, but developing a strong quantization recipe requires decisions about methods, data, and experimental evidence. We study frontier coding agents for this task. When given a quantization method and its target configuration, a strong standalone agent can set up the environment, debug the implementation, and reproduce the reported results. In this setting, however, the researcher still chooses the method. Autonomous quantization optimization instead requires the agent to select methods, design experiments, and decide which directions to pursue. Although the agent can develop competitive recipes, we observe four problems during autonomous exploration: overlooked prior work, inadequate experimental design, unsupported conclusions, and premature stopping. Based on these observations, we propose an Explorer–Reviewer system. The Explorer develops and executes candidate improvements, while an independent Reviewer evaluates research decisions at four checkpoints: proposals, experimental plans, conclusions, and requests to stop. The agents communicate through a shared research log. Our study covers 2-bit weight-only quantization of Llama-3-8B, Qwen3-8B, and Mistral-7B. On Llama-3-8B, Qwen3-8B, and Mistral-7B, our system improves average zero-shot accuracy across seven tasks by 0.88, 0.60, and 0.19 percentage points, respectively, over the standalone agent. It also outperforms the evaluated published quantization baselines on this metric for all three models. Case studies show how the Reviewer guides data selection, requests stronger experimental controls, and challenges unsupported conclusions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.