DiaKV: Offline KV Cache Training for Context-Level Adaptation
Abstract
Test-time training is a learning paradigm that updates the weights of the model on a test-time input, enabling more granular adaptation than fine-tuning. However, adaptation performed after a request arrives adds online training latency, limiting its use in latency-sensitive serving. We propose DiaKV, which trains the KV cache of a reused context offline before future queries arrive. DiaKV directly optimizes the KV cache of a context while keeping the model weights and cache length fixed. We formulate this adaptation as KV-cache posterior maximization given context-specific question-answer distribution, and augment 20 seed QAs per context into a larger training set. Across four long-context QA benchmarks and three 3–4B LLMs, DiaKV improves the primary macro metric in 11 of 12 model–benchmark pairs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.