acceptodds
Under review as a conference paper at ICLR 2027

What Rank Calibration Actually Measures: The Sparse, Repeatable, Query-Specific Response of Video LLMs to KV-Cache Compression

Abstract

Low-rank compression of a video language model’s key cache is usually tuned by reading a rank budget off the cache’s spectrum and reporting the change in average accuracy. We treat compression as a controlled intervention instead and ask where it acts. Holding the input and the query fixed and moving only the rank, the set D of inputs whose answer changes has three properties, measured on MVBench with Qwen2-VL-7B under rules frozen before the data. It is sparse: a task’s accuracy-versus-rank knee is carried by 5–14% of its inputs, which is why a per-task allocator gains little over the best uniform rank (+1.3pp, CI [−0.3, +3.0]). Its allocation-relevant part is unreadable from the signals we test: whether a cheap rank breaks an input is predictable (AUC 0.755), but whether more rank fixes it is at chance (AUC 0.47, [0.38, 0.57]). And it repeats: across option order, frame sampling, a second rank pair, real video, a second compressor and a second architecture, an input that changed its answer once changes it again with probability 56% against a base rate of 12% (n=2,392, κ=0.50). That repeatability belongs to the video–question pair: profiling a video on some questions does not inform its others (+0.19pp against a video-level oracle of +8.3). Together these identify what a rank calibration has to measure—the pair, not the task, the spectrum or the video—and imply a measure-then-reuse pattern that labels only the inputs where two ranks disagree (25–49% of the pool, an exact bound) and returns +2.4pp (CI [+1.25, +3.72]) on held-out realizations at a matched key-cache budget, in workloads whose outputs must be recomputed rather than looked up.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.