acceptodds
Under review as a conference paper at ICLR 2027

Fast and Accurate: Improving Personalized LLM Performance by Semantic Concentration Inspired Workload Split between RAG and Fine-Tuning

Abstract

With ever increasing personalized large language model (LLM) services, the next question to answer is how to improve the quality of services/users' experiences, which is indicated by personalized LLM generation accuracy/quality and inference latency. Although state-of-the-art designs, e.g., parameter-efficient fine-tuning (PEFT), retrieval-augmented generation (RAG) or Full PEFT+RAG, can help, they largely ignore the heterogeneous semantic concentrations of users' historical data and cannot fully unleash the potential of personal data for personalized LLM performance enhancement. In this paper, we propose PEFT Union RAG (PUR), a novel "fast and accurate” approach that exploits the semantic concentration of users' personal historical data to split workload between PEFT and RAG, aiming to improve personalized LLM performance in terms of better generation accuracy/quality and lower inference latency. Briefly, PUR partitions user's personal historical data and assigns semantic concentrated "core” information to PEFT for offline parametric fine-tuning, while retaining dispersed information for RAG. During the online LLM inference, PUR directly generates responses for "core” queries covered by the PEFT region and invokes RAG only for sporadic "peripheral” queries. Experimental results show that, by appropriately splitting workload between PEFT and RAG, PUR consistently achieves the best generation accuracy/quality for various users across multiple tasks, while reducing inference latency by 14.4%–46.6% compared with the best performed method.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.