acceptodds
Under review as a conference paper at ICLR 2027

Stable Curves, Unstable Items: Item-Level Scaling Heterogeneity in Video LLMs

Abstract

More visual input often improves average accuracy in video language models. However, aggregate gains do not reveal whether models preserve answers they already get right. We present a paired evaluation protocol that tracks the same questions across visual budgets within frozen models. Our study covers four models and four multiple-choice benchmark splits on a common budget grid, including text-only input. On these grids, – of questions exhibit at least one correct-to-incorrect transition as the budget increases. Choosing the best budget for each question in hindsight yields – percentage points higher accuracy than the best fixed budget. This gap reveals complementary successes that no single budget captures. We release over 120,000 per-item evaluation records and analysis code to support reproducible evaluation of visual scaling. Our findings show that larger visual budgets change which questions are answered correctly, not simply how many.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.