acceptodds
Under review as a conference paper at ICLR 2027

OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation

Abstract

Omni reference-to-video (R2V) generation is emerging as a general paradigm that enables flexible video customization from diverse and compositional visual references. However, existing benchmarks fall short of these emerging capabilities: their test cases cover only a limited range of reference types and compositions, and their evaluation protocols largely assess holistic reference consistency, overlooking whether reference factors are properly preserved, disentangled, and routed. Meanwhile, the high cost of constructing omni R2V training data makes suitable training resources scarce. To address these gaps, we introduce OmniVBench and the Omni-R2V Dataset, establishing a shared foundation for evaluating and training omni R2V models. OmniVBench expands R2V evaluation across broader reference types, richer reference compositions, and more flexible instruction-driven customization, covering 7 task families and 18 fine-grained tasks spanning content, motion, style, structure, narrative, and multi-reference settings. Complementing this task design, we introduce factor-grounded evaluation with 12,172 case-specific checklist items, explicitly assessing whether intended reference factors are faithfully preserved, correctly disentangled and bound to their targets, and realized according to the instruction. We further introduce the Omni-R2V Dataset, bringing industrial-grade training resources for diverse R2V tasks to the broader research community. Drawing primarily on a large-scale corpus of professional video footage, it comprises 340K processed training samples spanning diverse reference types and multi-reference compositions. Specifically, we develop task-specific pipelines for reference-target pair construction, offering a practical and scalable recipe for omni R2V data construction. Extensive evaluation of advanced open- and closed-source R2V models reveals substantial performance variation across diverse tasks and distinct capability patterns across evaluation dimensions, highlighting OmniVBench’s fine-grained evaluation capability.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.