KAP: Bridging the Knowledge Selection–Runtime Consumption Gap in LLM Systems
Abstract
Modern LLM systems increasingly rely on knowledge-selection pipelines that produce structured signals, such as evidence rankings, graph relations, multimodal alignment, and confidence scores. However, these signals are typically discarded when knowledge is serialized into a flat prompt, leaving the serving system unaware of which knowledge is important during decoding. This creates a gap between knowledge selection and runtime consumption: the model is presented with a complete logical context, while the runtime lacks the information needed to avoid unnecessary KV access. We propose Knowledge Access Planning (KAP), an execution abstraction that bridges this gap by carrying structured knowledge priors into LLM serving through a universal intermediate representation (IR), the runtime access plan. KAP separates logical prompt semantics from physical execution: its compiler translates knowledge-selection priorities into backend-executable plans, while its executor selectively accesses physical KV states according to these plans, without modifying the model, prompt semantics, or training procedure. We instantiate KAP with GraphSpec, a reference realization that executes runtime access plans within vLLM, and derive a conditional cost model characterizing its positive-speedup regime. Across a range of experimental settings, GraphSpec reduces proposal-path KV exposure to 5.1–7.9% of the source state at 128K and translates this reduction into long-context decode speedups while maintaining answer quality comparable to full-context decoding. These results demonstrate KAP's effectiveness in substantially reducing proposal-path KV exposure. More broadly, they show that knowledge-selection signals can be compiled into actionable execution plans, extending their role beyond prompt construction to efficient LLM decoding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.