acceptodds
Under review as a conference paper at ICLR 2027

Most of the Gain Is Not Video: Taking Apart a Fine-Tuned Video Question-Answering Score

Abstract

Video question-answering benchmarks rank video-language models, and fine-tuning on a benchmark's own training questions can lift a model's score by tens of points, a lift naturally read as better use of the video. A multiple-choice score can also rise for reasons unrelated to the video: a fine-tuned model learns the scorer's answer format, the answer statistics of the question set, and even answer letters the prompt never offers. Each effect is known in isolation, but their shares of a fine-tuned gain are rarely measured, and a leaderboard number cannot separate them. We measure them on UrbanVideo-Bench, where a LoRA adapter on Qwen2.5-VL-7B raises accuracy on a clip-disjoint development split from 29 to 71. With prompt, scorer and split fixed and every comparison paired on the same predictions, a fifth to a quarter of the gain is the benchmark's released scoring rule misreading the zero-shot model's replies, and a seventh comes from questions whose answer letter the prompt never offers. An adapter trained on the question text alone, never seeing a frame, recovers 89 to 91 of the gain and scores within three points of each of three video-trained seeds; permuting the answer options costs the video-trained reference six points. Two further readings of the score shrink under controls: an apparent 20-point simulated-versus-real gap at least halves once both simulators are pooled, and three re-trained seeds place the 71 at the top of its recipe's range. These checks are cheap; we collect them in a reporting checklist for fine-tuned video-QA scores and provide the code and aggregate data.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.