acceptodds
Under review as a conference paper at ICLR 2027

TickTock: Teaching and Evaluating Video LLMs to Tell Time at Frame-Level Precision

Abstract

Recent advanced video large language models show temporal awareness: they can relate events to when they occur and answer multi-hop temporal questions. Despite this potential, they remain far from frame-level precision, especially for complex events whose boundaries unfold within a fraction of a second. A key obstacle is that their training labels and benchmark references are mostly recorded in whole seconds, so neither supervision nor evaluation can tell whether a prediction points to the right frame. In this work, we introduce TickTock, a frame-anchored annotation framework for teaching and evaluating frame-level temporal grounding. Instead of guessing a timestamp, the annotator, either a human or advanced MLLMs, selects the boundary frame from a numbered filmstrip, and returns its exact timestamp, so every label is traceable to a real frame. Building on this framework, we construct TickTock-5K, a training set of 5,000 videos spanning 11 domains from short videos and sports to lectures and screen tutorials, and release refined annotations for the TimeLens-100K and TimeLens2-93K training corpora. For evaluation, we build TickTock-Refined, a frame-level re-annotation from TimeLens-Bench and VUE-TR, complemented by TickTock-Point500, 500 human-annotated moments snapped to decoded frames. Across 28 open and closed video-LLMs, the strongest model places 79.2% of moments within five seconds but no model exceeds 21.6% within 0.1 s, a gap that IoU-based metrics largely conceal. Frame-anchored references reorder models and track human point accuracy more closely. Together, TickTock and its data and evaluation protocol lay a foundation for frame-level temporal grounding in Video-LLMs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.