Frameworkers: A Dynamic Multi-Agent Framework for AI Video Production
Abstract
Modern video generators produce high-quality individual clips, but complete video production requires coordinating many interdependent steps, from scripting and storyboarding to generation and editing, together with the assets they pro duce and consume. Existing automated systems rely on fixed pipelines that adapt poorly to diverse inputs,while general-purpose large language models (LLMs)re main unreliable for long-horizon orchestration and multimodal asset routing. We introduce FRAMEWORKERS, which replaces the fixed production pipeline with learned, descriptor-conditioned orchestration over modular capabilities, grounded in a persistent Workspace. The Director treats production as dynamic task man agement, editing a Dynamic Task Stack to decide what to do next, which capabil ity to invoke, and how to replan the remaining work after feedback or failure. The Assistant executes each task against shared state that carries scripts, assets, and memory across steps. Capabilities are registered as sub-agents described only by natural-language descriptors, so adding one requires no change to the framework. We train the Director for this routing task with supervised fine-tuning followed by Group Relative Policy Optimization (GRPO). FRAMEWORKERS reaches 88.4% chain-level routing accuracy against 71.2% for the strongest LLM planner we evaluate,recovers from 96.7% of injected runtime failures with a general-purpose LLM as replanner and no recovery-specific training, correctly invokes sub-agents unseen in training, and is rated above single-agent and prior multi-agent video systems in automatic and human evaluation, with the largest human-rated gains on controllability and cross-shot consistency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.