acceptodds
Under review as a conference paper at ICLR 2027

TESSA: RETHINKING LORA FOR EFFICIENT MULTI- ADAPTER SERVING

Abstract

Low-Rank Adaptation (LoRA) enables parameter-efficient fine-tuning by learning small, task-specific updates. Serving requests with different LoRA adapters on a shared base LLM requires executing adapter-specific low-rank computations during inference while reusing the base model, enabling inexpensive adapter swapping and heterogeneous-adapter batching. While these computations are theoretically lightweight, multi-LoRA serving can incur substantial practical overhead due to two sequential, GPU-unfriendly matrix multiplications and additional inter-GPU communication under tensor parallelism. Specifically, despite adding less than 1% to both computation and parameter memory, LoRA adapters can reduce serving throughput by 15.5% with the reduction growing to 42.9% as tensor parallelism scales. Existing multi-LoRA systems, such as S-LoRA and Punica, optimize adapter execution but retain the two-stage shrink–expand structure, leaving GPU-inefficient small sequential matrix products and adapter-specific inter-GPU communication as persistent bottlenecks. In this paper, we propose Tensor-parallel Shard-aligned Slice Adaptation (TESSA), a new adapter structure tailored to efficient multi-adapter serving, rather than optimizing execution around the existing LoRA structure. TESSA trains contiguous, shard-aligned subsets of the base weight matrix and represents each adapter as a weight delta over the selected subset. Contiguous partial updates enable base and adapter computations to be executed as a single wider, GPU-friendly matrix multiplication. Moreover, shard-aligned placement eliminates additional adapter-specific communication under tensor parallelism. We demonstrate that TESSA substantially improves serving efficiency across diverse workloads while maintaining adaptation quality comparable to LoRA under approximately matched adapter parameter budgets. Our vLLM implementation achieves up to 3.46 the throughput of S-LoRA at approximately matched adapter budgets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.