Neural Context Compiler: Learning KV Caches for Long-Context LLMs
Abstract
Documents that are queried many times can be compiled once into a compact key–value (KV) cache that a frozen language model reads at every query. We show that such compaction must alter attention at almost every query and that a training objective by itself bounds the loss only on the inputs it supervises. The leading learned method, self-study, supervises answers to synthetic questions that must first be generated for each document, and hence only the facts those answers state. We introduce Neural Context Compiler (NCC), whose refinement supervises a prediction of every document token and generates no text: starting from an Attention Matching (AM) cache, it refines only the compact keys and values by distilling the full-context teacher on teacher-forced continuations of the document. With the hybrid Qwen3.6-27B on 64 held-out papers, a 512-row NCC cache raises constrained recall of masked source numbers from 62.3% to 79.8% over AM and Qasper F1 from 38.1 to 44.2, 28–44 points above StreamingLLM, SnapKV and KVzip. It is statistically indistinguishable from self-study on constrained recall and QA, 11 points better under greedy decoding and 13–25 points better on numbers self-study’s answers never state, with 92% less refinement time. On three of four LongBench QA tasks it matches the full-context reader with a 5–11×smaller state, and on 34K–97K-token documents at 11–30×compression it retains 74–84% of full-context QA F1 (AM: 19–49%). A fully generation-free variant does as well, an FP4 reader with a 4-bit state keeps its accuracy with a 14×smaller stored state, and the recall gain replicates on three further hybrids, fading where the teacher misses many numbers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.