acceptodds
Under review as a conference paper at ICLR 2027

RoboChrono: A Real Robot Benchmark for Streaming Task Understanding

Abstract

Existing video benchmarks mainly use human centric or passively recorded footage, offering limited direct evaluation of temporal understanding in robot task execution. We introduce RoboChrono, a benchmark with 39 scenarios and 34,754 distinct questions from controlled executions on GIMME_ARM and Marvin Pro, together with bare hand human recordings of five of the same tasks. RoboChrono covers diverse embodiments, end effectors, and viewpoints and evaluates seven capabilities grouped into recognition, alignment, and temporal grounding. Separate scores assess action understanding and the ability to connect observations across time and camera views. Evaluation of 17 vision language models reveals stable overall rankings across pairs of evaluation sets with no shared questions (Spearman –), yet aggregate scores conceal marked differences between capabilities. Doubao-Seed-2.0-Lite achieves the highest observed recognition accuracy (71.7%) and temporal grounding score (0.576 [email protected]). Gemini-3.6-Flash ranks first in alignment (cross view matching and frame ordering) but last in recognition, while placing thirteenth overall. Median frame ordering accuracy across models and sets is 28.9%, close to the 25% chance baseline. These findings expose gaps in temporal and cross view understanding and motivate capability specific evaluation alongside aggregate scores.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.