EventTimeQA: Auditing Intervention-Consistent Millisecond Temporal Representations
Abstract
Event-stream and high-frame-rate benchmarks usually score samples independently, which cannot show whether a model changes its prediction for the right reason when one temporal variable changes. We present EventTimeQA, a controlled synthetic evaluation of intervention-consistent millisecond temporal perception and representation. Its 10,000 factual/counterfactual pairs hold scene nuisance variables fixed while swapping onset order, relative frequency, or time-to-contact. We evaluate predictions with pair accuracy and flip consistency, temporal-resolution sweeps, timestamp perturbations, scene and sensor shifts, and an explicit audit of post-simulation event-retention policies. Frozen v3 evidence shows that lightweight supervised event representations approach ceiling performance, while visual-language performance depends strongly on interface: most tested event-contact-sheet interfaces remain weak, whereas a budget-compliant third-party Gemini route reaches pair accuracy on RGB500 but at most on event input. A leakage-audited 0/2/4/8-shot study further shows non-monotonic, model- and interface-specific effects. We release executable construction, shortcut, statistical, documentation, and maintenance protocols. The benchmark and all empirical claims are deliberately restricted to the controlled synthetic domain; physical-sensor transfer and human performance are outside the scope of this release.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.