LoRA-Ring: Locality-Preserving Consistent Hashing for Dynamic Adapter Routing in Large Language Models
Abstract
Serving large language models with many LoRA adapters requires routing each request to a server that has the relevant adapter loaded, since adapter loading latency (200–800ms) dominates request latency when adapters are not cached. Existing routers assign adapters to servers using static hash partitioning, which ignores adapter co-access locality—pairs of adapters that tend to appear in consecutive requests from the same user. LoRA-Ring introduces locality-preserving consistent hashing that places co-accessed adapters on adjacent ring segments, enabling a single server to serve common adapter pairs without cross-server fetches. We prove that LoRA-Ring reduces cross-server fetches by a factor proportional to the locality score of the workload, and show empirically that it reduces median request latency by 31% on a trace from a production multi-tenant serving system.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.