VideoForge: Toward Recursive Self-Improvement Environments for Agentic Video Evaluation
Abstract
Improving video evaluation agents requires identifying why they miss or misdescribe defects and determining which harness changes correct these failures. Existing methods typically optimize the evaluation harness with a fixed outer-loop procedure, leaving the improvement process itself unchanged. We introduce VideoForge, a recursive self-improvement environment that adapts both the video-evaluation harness and the procedure used to improve it. VideoForge exposes task execution, evidence inspection, harness revision, and self-reflection as interaction primitives, allowing improvement trajectories to adapt to observed failures and accumulated experience. To evaluate these capabilities, we introduce VidDiag, a benchmark comprising 1,087 real-world video generation and editing cases with 2,932 human-annotated defects spanning seven dimensions and 27 fine-grained failure types. The benchmark evaluates both defect-category identification and instance-level defect descriptions. Across two model–agent configurations, VideoForge improves overall category F1 by 3.6%–4.9% and strict semantic matching by 2.9%–4.6% over the corresponding initial harnesses, achieving the highest overall scores among the compared optimization methods, including SkillOpt and AHE. Ablations under the same task-execution budget further support the benefit of adapting the improvement procedure. We further conduct a separate human-evaluated video-debugging study on 100 cases to examine whether accurate diagnosis can support downstream repair. Providing ground-truth defect descriptions improves repair success by 11%–12%, yet the highest success rate remains only 32%. These results show that adapting the improvement process can strengthen fine-grained video diagnosis, while translating accurate diagnosis into effective repair remains challenging.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.