acceptodds
Under review as a conference paper at ICLR 2027

Quantifying GPU Energy Savings from Latent Communication in Multi-Agent LLM Inference

Abstract

Latent-space communication replaces decoded text between agents with hidden states and KV-cache transfer. Although prior work reports speed and token reductions, its energy implications have not been measured directly. We compare single-agent, text-based multi-agent (TextMAS), and latent multi-agent (LatentMAS) inference on identical NVIDIA A100 hardware for Qwen3-8B and LLaMA 3.1-8B-Instruct on GSM8K. Relative to TextMAS, latent communication reduces net GPU energy by 70.7% for Qwen3-8B and 80.6% for LLaMA. Text and latent communication differ by only 0.3 and 2.5 percentage points, respectively; question-clustered confidence intervals include zero and do not establish equivalence. The saving does not come from a cheaper individual latent step: latent steps and decoded tokens take similar time per unit (38.5 vs. 37.1 ms). Instead, the reduction is primarily associated with replacing hundreds of decoded intermediate tokens with 40 bounded latent steps. The single-agent comparison is model-dependent: for Qwen3-8B, latent communication is slightly more accurate, uses 16.6% less energy, and runs 1.18 faster than the matched single-agent baseline, whereas for LLaMA the single-agent baseline is both more accurate and lower-energy than either multi-agent mode. A second Qwen3 benchmark shows the same qualitative energy and accuracy pattern. These results quantify a substantial GPU-energy reduction from latent communication in the single-stream setting, while highlighting reduced transparency as an important practical limitation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.