Combining Document LoRA and KV Cache for Long-Context Question Answering
Abstract
Document LoRA encodes information from a source document through low-rank adaptation (LoRA), which trains small weight-update matrices while keeping the base model fixed. These document-specific parameters provide a form of memory alongside the key-value (KV) cache. We observe that the KV cache and document LoRA differ in the types of questions they answer correctly, motivating their combination for long-context question answering. Our approach first applies KV cache eviction, then fine-tunes the model on synthetic question-answer pairs from the source document using LoRA. Training updates only the low-rank matrices in the feed-forward projections while keeping the compressed KV cache in context. When answering a question, the model uses both the document LoRA and the compressed KV cache. On 2WikiMultihopQA, answering with both memories outperforms using the selected KV cache alone or the same learned adapter alone across all four models. The overall F1 gains reach 25.0 and 9.0 points, respectively. We expect this approach to improve efficiency for repeated QA over a shared document and, with chunked processing, for documents that exceed the model's context window.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.