acceptodds
Under review as a conference paper at ICLR 2027

AgentDirector: A Movie Recap Benchmark for Long-Horizon Multimodal Agents

Abstract

Existing multimodal benchmarks probe perception, reasoning, and generation in isolation, yet provide limited diagnostic insight into how these capabilities remain coordinated across long sequences of planning, tool use, and verification. We introduce AgentDirector, a 70-task benchmark that evaluates this capability through movie recap production from full-length films. The workflow couples narrative design and footage localization with narration synthesis, audiovisual editing, and inspection of intermediate results. We decompose this workflow into five settings spanning Planning, Execution, and End-to-End Production, varying supplied narration and plans, temporal requirements, and deliverables to examine these capabilities and their coordination. Our evaluation decomposes performance into Completeness, Verification, and Quality, with criteria aligned to each setting's responsibilities. These complementary components distinguish delivery failures, constraint violations, and weaknesses in narrative or audiovisual quality. The benchmark further supports systematic comparisons across agents and harnesses in task performance, editing quality, and computational cost. Component scores and execution traces help identify capability profiles and failure patterns throughout recap production.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.