acceptodds
Under review as a conference paper at ICLR 2027

LoRA-Ring: Locality-Preserving Consistent Hashing for Dynamic Adapter Routing in Large Language Models

Abstract

Serving large language models with many LoRA adapters requires routing each request to a server that has the relevant adapter loaded, since adapter loading latency (200–800ms) dominates request latency when adapters are not cached. Existing routers assign adapters to servers using static hash partitioning, which ignores adapter co-access locality—pairs of adapters that tend to appear in consecutive requests from the same user. LoRA-Ring introduces locality-preserving consistent hashing that places co-accessed adapters on adjacent ring segments, enabling a single server to serve common adapter pairs without cross-server fetches. We prove that LoRA-Ring reduces cross-server fetches by a factor proportional to the locality score of the workload, and show empirically that it reduces median request latency by 31% on a trace from a production multi-tenant serving system.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.