acceptodds
Under review as a conference paper at ICLR 2027

EvoVE: Evolving Video Evaluation Through A Self-Improving Harness

Abstract

Modern video generation models are improving at a rapid rate, and they are quickly spreading into domains such as film production, advertising, education, and robotics simulation. Each new generation brings new capabilities and, with them, new failure modes: Seedance 2.5 accepts new audio references and emits longer clips than Seedance 2.0, and its most consequential errors (e.g., Audio-Visual Speech Misalignment) simply did not exist for its predecessor. Benchmarks, however, do not move with the models. Therefore, we introduce EvoVE (Evolving Video Evaluation), a model-agnostic, recursively self-improving evaluation harness for AI-generated video and multimodal understanding. EvoVE jointly evolves an evaluation plan, a trace-analysis textbook that serves as memory, and a dynamic tool bank. On error taxonomy benchmark, EvoVE achieves a macro-F1 of 88% at 5-shot, compared to the 67% VLM baseline. We formulate harness evolution as meta-learning, enabling harness self-improving. With only 5-shot target examples, EvoVE self-improves by 11.9% F1 score in-distribution and 7.8% F1 out-of-distribution. Finally, with only a task-adapter change, EvoVE generalizes from video evaluation to image benchmarks, achieving 82%–93% held-out accuracy on TextVQA, MMBench, and MMMU.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.