SOMA: Efficient Multi-turn LLM Serving via Small Language Model
Abstract
Large Language Models (LLMs) are increasingly deployed in multi-turn dialogue settings where preserving conversational context across turns is essential. A standard serving practice concatenates the full dialogue history at every turn, which reliably maintains coherence but incurs substantial cost in latency, memory, and API expenditure, especially when queries are routed to large proprietary models. Existing approaches often struggle to balance the trade-off between response quality and efficiency. We propose a framework that exploits the early turns of a session to estimate a local response manifold and then adapt a smaller surrogate model to this local region for the remainder of the conversation. Concretely, we train soft prompts that push the small surrogate away from the large model's responses, with an anti-degeneration term for stability, and score each early turn by how much agreement survives. Localized LoRA fine-tuning then distills the turns with robust agreement, so the surrogate runs without prompts at inference. A gate then hands the session to the surrogate and rolls back on drift. We further provide a theoretical analysis for key components in SOMA. Experiments on six multi-turn benchmarks suggest that SOMA largely preserves the original model's responses and lowers serving cost for medium-to-long sessions that stay on topic. The source code is provided at: https://anonymous.4open.science/r/SOMA-4175 .
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.