acceptodds
Under review as a conference paper at ICLR 2027

Keep Your Eye on the Cup: Diagnosing Identity and State Tracking in Video Language Models

Abstract

A wrong answer on a video benchmark can reflect missed motion, confused object identities, or an incorrectly recovered state. We examine these demands with a con- trolled shell game in which cups swap positions while carrying hidden balls. The benchmark contains 650 synthetic videos and 62 exploratory staged real videos. Its central experiment adds persistent identity markers to the same synthetic videos, holding motion, questions, and answers fixed. On monochrome videos, labeled boxes improve target-cup final-position accuracy by 10.6–65.4 percentage points across eight models. The equal-model gain is 37.5 points for final ball position but only 8.6 points for counting a ball’s visits to a position. Accuracy remains lower on longer sequences even with labels. In a separate reconstruction task, only 8.9–30.0% of videos have every reference state correct. Increasing the requested sampling rate from 2 to 4 FPS has mixed effects. These results expose sensitivity to identity information without establishing that every intervening swap was tracked: marked endpoint questions can be answered from initial and final observations

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.