Tracing Arrow-of-Time: Diagnosing and Addressing Temporal Information Loss in Video-LLMs
Abstract
The Arrow-of-Time (AoT) task requires distinguishing forward from backward video playback by recognizing temporally irreversible events. Prior work has revealed a substantial gap between humans and Video-LLMs on AoT in end-to-end evaluations. Here, we use AoT as a diagnostic lens to localize where this gap arises: do visual backbones fail to encode temporal information, or is such information lost when transferred to the LLM? We combine vision-encoder probing with downstream video question answering (VQA) to trace temporal information across the encoder, projector, and LLM. We find that video-centric encoders with explicit temporal modeling encode substantially stronger temporal signals than the evaluated frame-centric encoders. However, encoding alone is insufficient: projector design and token compression can disrupt temporal information before it reaches the LLM. Moreover, layer-wise analysis reveals the distinction between probe-level recoverability and LLM-level usability, underscoring the complementary roles of probing and VQA. Guided by these findings, we combine a temporally aware video encoder with a time-preserving projector and achieve 98.1% accuracy on AoT, surpassing human performance. Finally, we show that AoT supervision improves temporal reasoning across both short- and long-video benchmarks. Our results position AoT as both a diagnostic lens and an effective supervision for improving broader temporal reasoning in Video-LLMs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.