acceptodds
Under review as a conference paper at ICLR 2027

Your Filter Picks the Winner: Predictable Selection Bias in Difficulty-Filtered Evaluation of Visual Token Compression

Abstract

A visual-token-compression benchmark scores methods only on items its "difficulty filter" (downsampling) fails. Phang et al. (2022) found adversarially filtered rankings unstable; we derive the distortion. In a threshold model, one-dimensional difficulty makes filtering the rescaling : rankings above the filter's accuracy survive, its not unfair. With an operator-specific component it decreases strictly, at fixed accuracies, in the tetrachoric correlation of the two failure patterns: a filter measures dissimilarity from itself, invertibly given the marginals. On 1.1M generations from six VLMs, difficulty is not one-dimensional (–; – net of cell-level noise on the open models), and that correlation spreads by –, net of sampling noise, over a filter's competitors, which, by the theorem, can reorder them. The hard-subset winner departs from the unfiltered order in 11–17 of 45 cells per model against 6–8 by sampling, significantly on both LLaVA models (Qwen2-VL: ). On VTC-Bench's model and filter, its conclusion that designed methods beat random selection holds at four of five budgets, filtered or not; a compressor filter can reverse it. The bias is choosable: a method's most flattering filter, picked on half the items, lifts its held-out rank by +0.9 to +1.0 (open models). It buys nothing: a matched-size random subset tracks the unfiltered ranking better on all six models (– against –) and creates at most 1.5% unsupported orderings, inverting almost none, versus the failure set's 2–11%. Sample uniformly; for hard items, cross-fit a panel. Code, raw generations and analysis scripts will be released at https://github.com/xxx/xxx upon acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.