SLMCoach: Context-Adaptive LM-Head Caching with Cross-Turn Reuse for Efficient Small Language Model Inference
Abstract
Small language models (SLMs) have emerged as a promising solution for privacy-preserving, low-latency artificial intelligence on edge devices. However, SLM inference remains constrained by memory footprint and data-movement overhead, with the vocabulary-dependent language modeling head (LM-head) constituting a substantial bottleneck. Existing vocabulary-reduction methods reduce these costs but typically treat queries independently, overlooking the temporal locality inherent in multi-turn interactions and missing opportunities to reuse LM-head rows across turns. To address this limitation, we formulate the cross-turn LM-head caching problem and introduce SLMCoach, a context-adaptive framework for efficient multi-turn SLM inference. By exploiting contextual continuity across dialogue turns, SLMCoach selectively retains LM-head rows with high reuse potential and incrementally updates a budget-constrained cache, reducing redundant parameter movement across consecutive queries. Empirical evaluations on a multi-turn dialogue benchmark demonstrate that SLMCoach reduces the LM-head footprint by 99.2% and peak GPU memory by 51.5% relative to full-head inference while preserving response quality. At matched response fidelity, SLMCoach further reduces LM-head data movement by 7.5× compared with query-isolated loading. These results highlight the benefit of exploiting cross-turn reuse for resource-constrained SLM inference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.