GRACE: Graded Representation Augmentation for Cost-Effective SLM–LLM Collaboration
Abstract
Large language models (LLMs) provide strong generation capabilities but incur substantial inference cost, while small language models (SLMs) offer efficient decoding at the expense of reduced performance on complex tasks. Recent studies have begun to explore ways of combining their complementary strengths, including speculative decoding and model switching. These efforts show the potential of involving stronger LLMs to assist SLM generation, but such assistance may still incur considerable additional computation and decoding latency. We propose GRACE, a Graded Representation Augmentation framework for Cost-Effective SLM-LLM collaboration, which improves SLM generation quality with minor additional LLM computation. GRACE establishes asynchronous bidirectional latent interaction between the frozen SLM and LLM, enabling LLM assistance without interrupting the SLM's autoregressive decoding. During decoding, a context aligner incrementally compresses the evolving SLM state into compact latent tokens to synchronize with the LLM, while representation augmentation blocks extract latent task-relevant guidance from the asynchronously available LLM representations to augment the SLM. These augmentations are performed at depth-aligned computation grades, allowing subsequent SLM layers to progressively absorb the guidance. Extensive experiments on instruction-following and mathematical benchmarks show that GRACE consistently improves SLM generation quality with favorable performance-efficiency trade-offs across different model scales, architectures, and versions, and remains effective when the collaborating models come from different model families.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.