acceptodds
Under review as a conference paper at ICLR 2027

From Hours to Tasks: Evaluating Evidence-Grounded Decisions by Long-Video Agents

Abstract

Long-video agents must integrate observations across time, track changing objects and states, and use tools to act. Evaluation requires executable tasks with agent-accessible evidence for key choices, so missing evidence is not mistaken for limited capability. We introduce EvidBind-Bench: 3,000 tasks across 13 domains grounded in hour-scale videos, with a design–binding method that fixes goals, constraints, and required decisions before linking them to observations, objects, times, and operational conditions. The framework comprises (1) Automated task generation, which instantiates reusable decision structures, binds choices to evidence, and generates instructions and execution paths, revising unsupported choices under unchanged requirements;(2) Task review and repair, which combines automated checks and human review, reuses shared assessments, and targets affected decisions, with acceptance checking support, executability, and requirement preservation; and (3) Reference-agent evaluation, which supplies ReAct and multi-agent frameworks with common public interfaces, scoring, and diagnostic execution records. Independent human audits of 390 sampled tasks estimate collection validity. Under the evaluated conditions, the complete procedure lowers combined labor and computational cost per estimated valid task by 72.86% relative to per-task review and repair. ReAct and the multi-agent system achieve 11.23% and 17.73% success, versus 91.73% for humans, revealing persistent difficulties on tasks passing construction checks. EvidBind-Bench combines low-cost task construction and agent evaluation to study the gap between solvability and reliable execution.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.