acceptodds
Under review as a conference paper at ICLR 2027

Quantize the Team: Joint Quantization and Topology Adaptation for Multi-Agent LLM Systems

Abstract

Multi-agent LLM systems solve tasks by exchanging information across repeated model calls. The large number of calls and long text sequences make efficient model deployment particularly important, with post-training quantization (PTQ) serving as a key approach. However, preserving the model's performance on individual calls may not preserve the benefits of communication: a message that helped another agent before quantization may no longer improve the final answer, even though it is still delivered. We call this loss a contraction of the system's effective topology: the communication graph stays the same, but fewer connections contribute to task performance. We characterize two failure mechanisms: sending agents provide less information beyond what other agents already supply, or receiving agents make less use of useful messages. To address these failures, we propose Quantization-Aware Repair or Reroute (Q-RoR) for agents that share a low-bit model. First, we introduce a unified quantization objective that jointly preserves message generation and use. By bounding the change in each message's edge value, the drop in task score when that message is removed, it grounds their optimization in what messages contribute and avoids manually balancing separate surrogate losses. Second, we propose an activation quantization operator tailored to multi-agent collaboration: it jointly controls quantization errors across roles to preserve complementary information. We pair this operator with attention calibration to restore receivers' use of useful messages. Third, compatibility-guided rerouting goes beyond quantization on a fixed topology or topology search with a fixed model: it jointly adapts the communication topology and quantized model to restore useful communication when repair alone is insufficient. Experiments across three models and five tasks, plus two further precisions on the main model, show that Q-RoR outperforms PTQ and sequential topology-search baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.