acceptodds
Under review as a conference paper at ICLR 2027

What to Distill and What to Retrieve: Knowledge-Split Distillation for LLM–SLM Collaboration

Abstract

Large–small language model (LLM–SLM) collaboration enables SLMs to leverage the stronger knowledge and reasoning capabilities of LLMs while reducing dependence on repeated LLM assistance. Across different collaboration paradigms, LLM-provided information is ultimately either internalized into SLM parameters or retained externally for on-demand access. Yet existing approaches largely treat this placement as a consequence of the collaboration paradigm, leaving a fundamental question under-explored: what should the SLM learn, and what should it retrieve? We find that LLM-provided information can be broadly divided into domain-specific knowledge and generalizable reasoning capabilities, which favor different placement strategies. Domain-specific knowledge primarily serves particular domains or tasks. Parameterizing such knowledge in a capacity-limited SLM may introduce knowledge interference without necessarily improving its generalizable capabilities, making external storage more suitable. In contrast, generalizable reasoning capabilities, such as the ability to select, organize, and apply knowledge, can transfer across problems and tasks and are therefore better internalized into SLM parameters than repeatedly retrieved during inference. However,such differentiated placement is challenging because domain-specific knowledge and generalizable reasoning are often tightly coupled within the same reasoning trajectory, and their roles depend on the dependencies among reasoning steps. The key challenge is therefore to disentangle them based on these dependencies while preserving a complete reasoning path. We propose KSplit, a type-aware knowledge-placement framework for LLM–SLM collaboration. KSplit represents LLM reasoning as a typed dependency graph connecting knowledge units and reasoning operations through their functional dependencies. Dependency-Aware Knowledge Partitioning (DAKP) exploits these dependencies to disentangle domain-specific knowledge from generalizable reasoning while preserving complete reasoning paths. Functional Knowledge Placement (FKP) then internalizes generalizable reasoning through knowledge-free reasoning supervision, while compiling domain-specific knowledge into an external retrievable repository. Experiments across multiple benchmarks show that KSplit consistently improves LLM–SLM knowledge transfer, with controlled studies further validating its type-aware partitioning and placement strategy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.