acceptodds
Under review as a conference paper at ICLR 2027

When the Window Cannot See the Question: Token-Denominated Observation Windows Are Not Language-Neutral

Abstract

KV-cache evictors of the SnapKV family decide what to keep by scoring the cache from the last w tokens of the prompt alone. When those tokens do not hold the question, the ranking is blind: it keeps the question's own keys and drops most of the gold passage's. A padded instruction, a JSON schema and a tokenizer each put the question outside the window, and the tokenizer is the lever an English-validated benchmark cannot see. On Qwen3-4B one line of instructions and the chat suffix leave 25 tokens after the question in English, 107 in Bengali and 167 in Telugu, so a research library's default w=64 leaves English untouched and costs Bengali and Telugu double digits. The fix is a check, not a tuned constant: ŵ = c + Q90, the trailing block plus a question-sized slack, both measured before any forward pass, says whether a shipped window is safe for a prompt family, returns the smallest that is, and says when none is. It removes the blind mode and restores the passage, +14 over the default on the full Telugu pool at a tenth more prefill time; for every public static tail we found, a fixed 256-token window does as well. Neither removes the eviction cost a fully visible question still pays. Because no window costs English anything at the shipped ratio, an English-validated benchmark cannot tell the two shipped constants apart; the window is one token-denominated constant among several, and the check generalises: measure the constant on the tokenizer before trusting an English validation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.