KV Thickets: Task Experts from Random Prefix-Cache Perturbations
Abstract
Neural Thickets reveal task-improving experts through random weight perturba- tions, but each candidate defines a distinct parameter state. In-context learning provides a natural starting point: a prefix steers a frozen model, and its key–value (KV) states permit continuous changes without editing discrete tokens. Can mod- est populations of small random prefix-KV perturbations also contain useful task experts? We introduce KV Thickets, which independently sample perturbations from prespecified distributions, select candidates on labeled examples, and reuse the selected perturbations across inputs while keeping both model weights and prefix text unchanged. Our vLLM implementation batches their evaluation under shared weights. Across seven models and seven benchmarks, 32 candidates yield a selection-split improver in 45 of 49 settings; at 256 candidates, coverage reaches all 49, with a mean improving fraction of 28.7%. On examples disjoint from selection, the top selection-ranked candidate and an eight-candidate plurality vote improve mean task score by 2.70 and 4.83 points, respectively, over the unperturbed greedy baseline. These comparisons are descriptive because the evaluation includes previously examined benchmark pools. Weight-space experts provide a comple- mentary reference: their induced cache changes have mixed direct effects, but their blockwise magnitudes define another useful random-search profile. Together, these findings identify fixed-prefix KV states as a search space for reusable task adaptation under a single frozen model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.