SEA-Bench: Self-Evolving Agentic Benchmark for Video Generation
Abstract
Video generation models are advancing rapidly, while their evaluation benchmarks remain largely static. Existing benchmarks typically fix the evaluated capability dimensions and task distributions, causing previously informative tests to saturate and newly emerging capabilities to remain unmeasured. We introduce SEA-Bench, a self-evolving benchmark construction framework that enables agents to discover and operationalize evaluation dimensions from observed model behaviors. Starting from an existing benchmark, SEA-Bench iteratively proposes capability hypotheses, designs executable prompts and evaluation procedures, validates them on target models, and uses an evidence-based decision agent to accept benchmark updates supported by the resulting evaluation evidence. Unlike prior dynamic evaluation methods that adapt instances or evaluation trajectories within a predefined capability space, SEA-Bench evolves the evaluation space itself. We instantiate SEA-Bench for text-to-video generation and evolve benchmarks for model pools with different capability profiles. Retrospective comparisons with independently developed video benchmarks assess the community alignment of retained dimensions, while human studies validate their meaningfulness, prompt faithfulness, and agreement between executable evaluators and human preferences. We compare aggregate model discrimination before and after benchmark evolution in the two evaluated model pools.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.