acceptodds
Under review as a conference paper at ICLR 2027

BudgetKV: Pre-Attention KV Retention with Resident Input Embeddings

Abstract

Long-context inference on memory-constrained devices must preserve nonlocal evidence while the key–value (KV) cache grows throughout prompt prefill. A request can therefore exceed available memory even when the model weights and intended final cache fit. Block-wise prefill can bound persistent state, but retention policies may still require contextual attention profiles, patched-prompt forwards, or auxiliary semantic encoders before deciding what matters. We introduce BudgetKV, a training-free method for requests whose task query is known before shared-context prefill. BudgetKV scores prompt tokens against the target model's resident input embeddings before contextual attention statistics are available, then refines a frozen lexical anchor through guarded equal-size block exchanges. This reuse of model parameters provides a query-conditioned pre-attention signal without another Transformer forward or learned encoder; under protected-block feasibility, persistent-KV token residency is at most B after selection and B+p between selections. On a frozen 400-document LongBench study, resident-embedding refinement yields a statistically supported +6.81-point gain for Qwen3-1.7B at B=1024, while the other tested model–budget cells remain inconclusive after adjustment. At 32K, the largest observed advantage is on needle retrieval rather than reasoning. On an 8GB Jetson, BudgetKV completes all six tested 8K requests at both budgets, whereas Full KV completes two; relative to successful Full-KV runs, persistent-KV footprint falls by 85.9–92.2%. These results support resident input embeddings as a useful pre-attention signal when contextual profiling is constrained by prefill memory, but benefits remain model- and task-dependent and scheduling adds overhead.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.