Keep Your Eye on the Cup: Diagnosing Identity and State Tracking in Video Language Models
Abstract
A wrong answer on a video benchmark can reflect missed motion, confused object identities, or an incorrectly recovered state. We examine these demands with a con- trolled shell game in which cups swap positions while carrying hidden balls. The benchmark contains 650 synthetic videos and 62 exploratory staged real videos. Its central experiment adds persistent identity markers to the same synthetic videos, holding motion, questions, and answers fixed. On monochrome videos, labeled boxes improve target-cup final-position accuracy by 10.6–65.4 percentage points across eight models. The equal-model gain is 37.5 points for final ball position but only 8.6 points for counting a ball’s visits to a position. Accuracy remains lower on longer sequences even with labels. In a separate reconstruction task, only 8.9–30.0% of videos have every reference state correct. Increasing the requested sampling rate from 2 to 4 FPS has mixed effects. These results expose sensitivity to identity information without establishing that every intervening swap was tracked: marked endpoint questions can be answered from initial and final observations
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.