acceptodds
Under review as a conference paper at ICLR 2027

AIRMAS: Inter-Agent Intermediate Activation Reuse for LoRA-Based Multi-Agent System

Abstract

LLM-based multi-agent systems assign a different role to each agent. Fine-tuning a separate model for every role is expensive in both compute and storage. To reduce the burden, recent systems share one frozen base model and give each agent an additional role-specific LoRA adapter. When the agents run in sequence, each agent receives the context produced so far and appends its own output. Because every adapter transforms the same tokens differently, each agent prefills the shared context again and keeps its own KV cache. Prefill cost and memory therefore increase with the context length. To address this issue, recent work splits the value cache into two parts. All agents share the base cache computed with the frozen backbone weights, and each agent keeps a small low-rank part from its own LoRA adapter. Reusing the low-rank part across agents requires every adapter to share a single down-projection. Adapters with independently trained down-projections cannot reuse it. We propose AIRMAS, an inference framework that addresses this limitation. Instead of sharing a single down-projection matrix, AIRMAS maps the stored low-rank cache of an earlier agent to the current agent with an matrix precomputed from the two down-projections. On agentic question-answering benchmarks, AIRMAS improves accuracy by up to 3.33% over BaseLRShared, the low-rank cache sharing scheme of LRAgent. It requires no change to the adapters, and its time-to-first-token stays within 2.4% of that scheme.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.