AgentDirector: A Movie Recap Benchmark for Long-Horizon Multimodal Agents
Abstract
Existing multimodal benchmarks probe perception, reasoning, and generation in isolation, yet provide limited diagnostic insight into how these capabilities remain coordinated across long sequences of planning, tool use, and verification. We introduce AgentDirector, a 70-task benchmark that evaluates this capability through movie recap production from full-length films. The workflow couples narrative design and footage localization with narration synthesis, audiovisual editing, and inspection of intermediate results. We decompose this workflow into five settings spanning Planning, Execution, and End-to-End Production, varying supplied narration and plans, temporal requirements, and deliverables to examine these capabilities and their coordination. Our evaluation decomposes performance into Completeness, Verification, and Quality, with criteria aligned to each setting's responsibilities. These complementary components distinguish delivery failures, constraint violations, and weaknesses in narrative or audiovisual quality. The benchmark further supports systematic comparisons across agents and harnesses in task performance, editing quality, and computational cost. Component scores and execution traces help identify capability profiles and failure patterns throughout recap production.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.