AuLoRA: An Automatic, Label-Free Framework for KV Cache Sharing in Multi-LoRA Systems
Abstract
In modern multi-agent and multi-expert systems, KV heterogeneity across models prevents a successor from inheriting its predecessor's KV cache: every handoff re-prefills the entire prefix context, inflating time-to-first-token (TTFT). This paper targets KV cache sharing in multi-LoRA systems, where one base model serves many domain-specific LoRA experts—a memory-efficient multi-expert setup. Here the base model is the natural sender: its KV cache is adapter-agnostic and can be reused by every expert. Existing training-free sharing methods, however, target either a pair of fine-tuned models or agents collaborating on a single task, leaving the base-to-adapter direction largely unexamined. Reusing the base model's KV cache directly leads to substantial performance degradation. We decompose the difference between the base model's and the expert's KV caches into a weight term and a hidden-state term, which correspond to the two existing families of training-free remedies, a low-rank correction and partial recomputation. Experiments show that the two terms are nearly orthogonal, and the hidden-state term dominates, so an exact low-rank correction removes only 1–3% of the KV difference. Recomputation is thus unavoidable in multi-LoRA systems. State-of-the-art recomputation methods rely on offline calibration: they search over consecutive-layer combinations for the one with the best task accuracy on a labeled calibration set, at a cost that grows quadratically with model depth. We present AuLoRA, a training-free KV cache sharing framework for multi-LoRA systems with two components. (i) Two-stage profiling: a label-free algorithm that reduces the search cost complexity from to , where is the model depth, without per-candidate decoding and with few calibration samples; its profiling result matches the current state of the art, with about faster end-to-end profiling on question answering and over an order of magnitude faster on summarization. (ii) Tail compensation: a LoRA-only mechanism that lets each expert prefill the tail of its own question on top of the shared context KV cache. With negligible compute overhead, it improves quality at a fixed recomputation budget and lowers the budget required for a target quality. Overall, with roughly one third of the layers recomputed, the two components recover 53–95% of the quality lost to direct reuse and reduce the expert's prefill latency by –.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.