acceptodds
Under review as a conference paper at ICLR 2027

Diagnosing Temporal Instance-State Binding and Its Stage-Dependent Repairability in Vision-Language Models: A Controlled Study

Abstract

Temporal visual understanding requires more than detecting that a state change occurred. A model must also identify which persistent instance underwent that change. We introduce TIB-Bench, a controlled synthetic benchmark designed to isolate event recognition from instance-state binding. Its counterfactual pairs end in pixel-identical final frames, but differ in history and in the correct event owner. On Qwen3-VL-2B, perception and event occurrence are both correct on 210 of 480 episodes. Yet the owner is wrong in 109 of those cases (51.9%), revealing a failure that event-level accuracy does not capture. Layerwise probing shows that owner-related information remains accessible in downstream representations, while late-layer activation interventions perturb binding decisions substantially more than event judgments. We then examine how this failure can be repaired. Replacing matched ordinary examples with counterfactual-generated examples raises counterfactual-history consistency from 0.006 to 0.850 without an explicit pairing objective. Under matched eight-layer LoRA budgets, early and middle adaptation also substantially outperform late adaptation. Temporal-order perturbations further show that repaired models rely more strongly on the canonical history. Together, these results separate event recognition, owner-related representation, and history-sensitive owner attribution within a controlled temporal visual setting.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.