acceptodds
Under review as a conference paper at ICLR 2027

MuLA: One Prefill, Many Agents via Efficient KV Cache Reuse in Multi-LoRA Agents Systems

Abstract

The multi‑LoRA agent framework, in which a collection of specialized agents share a single base LLM while each is equipped with its own LoRA adapter, has been widely adopted in complex real‑world applications. In such pipelines, a downstream agent has to prefill the context produced by upstream agents before generating its own response. Because upstream agents typically accumulate long contexts, re‑prefilling the entire context from scratch for every downstream agent incurs prohibitive computational overhead and inference latency. Yet we find that this repeated prefill is largely redundant: the KV cache of upstream agents can be substantially reused, requiring only lightweight sparse corrections. To this end, we propose **MuLA** (Multi‑LoRA‑Agents Sparse KV Cache Migration), a framework that efficiently reuses upstream KV caches by selectively recomputing only a small subset of critical token positions. Our empirical analysis shows that LoRA‑switch‑induced KV‑cache deviations are sparse over tokens and delayed to deep layers. Critically, cheap attention‑concentration signals collected in the upstream agent’s forward pass strongly anticipate these error hotspots. Building on this insight, we exploit attention concentration as a cheap early indicator to proactively identify tokens requiring correction when switching LoRA adapters between agents, thereby completing sparse recomputation before the target deviations fully manifest, while balancing both accuracy and efficiency in the agents workflow. Results show that **MuLA** preserves nearly full‑prefill quality with a small fraction of token recomputation, delivering substantial downstream prefill speedup and promising end‑to‑end speedup.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.