RobotEQ-Video: A Video-Centric Benchmark for Social Proactive Intelligence with World-State Taxonomy
Abstract
Social Proactive Intelligence (SPI) extends proactive assistance beyond task completion to consider social appropriateness in diverse embodied scenarios. However, prior SPI research faces two key limitations. First, existing work focuses on static images, whereas dynamic videos provide crucial cues for inferring human states and needs, offering richer information than isolated images. Second, prior work often relies on free-form data collection pipelines, which fail to guarantee comprehensive coverage of diverse contextual factors. To address these gaps, we introduce **RobotEQ-Video**, shifting the focus from *image-centric* to *video-centric* analysis. To ensure coverage of diverse contextual factors, we construct a *hierarchical world-state taxonomy* organized into a four-level coarse-to-fine structure, comprising 6 domains, 20 dimensions, 142 level-1 attributes, and 816 level-2 attributes. The resulting benchmark comprises 2K+ videos with 100K+ human annotations and 17K+ labels for assessing action appropriateness. Benchmark evaluation reveals that current systems still fall short of human performance. We further explore how world models can help tackle this task. *This work advances SPI research from static images to dynamic videos and ensures more comprehensive coverage of diverse contextual factors through our world-state taxonomy.*
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.