InvisibleLoRA: Fast Multi-LoRA Serving Without Per-Token LoRA Decoding
Abstract
Low-rank adaptation (LoRA) specializes a single base model into many fine-tuned variants, and production systems now serve hundreds of adapters on a shared base model. Serving many adapters concurrently, however, introduces substantial inference overhead. Because requests are partitioned across active adapters, each low-rank computation covers only a small fraction of the batch and runs as a small, GPU-inefficient operation. Consequently, per-step adapter overhead grows with the number of active adapters even though each adapter holds only a small fraction of the base-model parameters. Existing multi-LoRA systems reduce the cost of each adapter execution through specialized kernels and scheduling, but every active adapter still executes at every token. We propose InvisibleLoRA, a speculative decoding method that uses the shared base model as the common drafter for every served adapter. A LoRA-off forward pass proposes the next token for every request in the batch, and a single LoRA-on forward pass verifies all proposals. InvisibleLoRA therefore applies adapter computation once per decoding cycle rather than once per token, while speculative verification preserves the output distribution of standard multi-LoRA decoding. InvisibleLoRA further introduces a hybrid KV cache: the LoRA-off base model writes speculative KV entries into the KV cache of the LoRA-on target model, and the target model overwrites the speculative entries during verification, so the drafter needs no separate KV cache and reads LoRA-conditioned KV states for the committed prefix. InvisibleLoRA also adopts a run-time gate that skips drafting on steps where the adapter cost cannot repay a draft pass. InvisibleLoRA reaches median draft acceptance rates of 92–97% across 456 publicly available LoRA adapters on five base models. Implemented in vLLM, InvisibleLoRA delivers lossless acceleration over an optimized multi-LoRA baseline, improving throughput by up to 1.84× and reducing per-token latency by up to 48%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.