acceptodds
Under review as a conference paper at ICLR 2027

SLMCoach: Context-Adaptive LM-Head Caching with Cross-Turn Reuse for Efficient Small Language Model Inference

Abstract

Small language models (SLMs) have emerged as a promising solution for privacy-preserving, low-latency artificial intelligence on edge devices. However, SLM inference remains constrained by memory footprint and data-movement overhead, with the vocabulary-dependent language modeling head (LM-head) constituting a substantial bottleneck. Existing vocabulary-reduction methods reduce these costs but typically treat queries independently, overlooking the temporal locality inherent in multi-turn interactions and missing opportunities to reuse LM-head rows across turns. To address this limitation, we formulate the cross-turn LM-head caching problem and introduce SLMCoach, a context-adaptive framework for efficient multi-turn SLM inference. By exploiting contextual continuity across dialogue turns, SLMCoach selectively retains LM-head rows with high reuse potential and incrementally updates a budget-constrained cache, reducing redundant parameter movement across consecutive queries. Empirical evaluations on a multi-turn dialogue benchmark demonstrate that SLMCoach reduces the LM-head footprint by 99.2% and peak GPU memory by 51.5% relative to full-head inference while preserving response quality. At matched response fidelity, SLMCoach further reduces LM-head data movement by 7.5× compared with query-isolated loading. These results highlight the benefit of exploiting cross-turn reuse for resource-constrained SLM inference.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.