ProAgent: High-Fidelity Product Video Generation via a Multi-Agent Framework
Abstract
Recent text-to-video models have made advertising video generation attractive as a way to cut production time and cost, yet direct prompting often yields videos that drift from product identity, contain hallucinations, distort printed text, and fail to convey a clear commercial expression. We present ProAgent, the first to address high-fidelity product video generation—a training-free multi-agent framework that turns a user prompt identifying a specific product into an advertisement-ready video end-to-end. ProAgent grounds this intent in retrieved product evidence, plans a shot-level blueprint and rewrites it into a generation prompt under explicit product constraints, and iteratively improves the result through a closed-loop critique-and-recovery process. To evaluate the quality of generated advertising videos, we further introduce PV-Bench (Product Video Benchmark), a multi-dimensional benchmark scoring videos across six product-centric dimensions. Experiments across 12 product categories spanning major e-commerce industries and multiple state-of-the-art video generators show that ProAgent consistently outperforms baselines on both PV-Bench and general metrics for video quality and semantic fidelity.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.