Are Image-to-Video Generation Models Robust to 3D Visual Illusions?
Abstract
Image-to-video generation models have shown strong ability to produce visually realistic videos with coherent motion and dynamic scene evolution from a single input image. However, their robustness under real-world 3D visual illusions remains unclear: in such cases, the first-frame appearance may conflict with the observable spatial, optical, projection, surface, object-identity, or lighting relations. Camera or object motion should remain compatible with the observable geometric, optical, and structural constraints, rather than preserving the illusion. Existing video generation benchmarks do not systematically evaluate failures in this illusion setting. To systematically evaluate this underexplored setting, we introduce VIVG-Bench (Visual Illusion Video Generation Benchmark), a benchmark of 680 carefully curated image-to-video prompts spanning diverse real-world 3D visual illusions. Each prompt is written without illusion descriptions: it removes explicit descriptions of the illusion mechanism and specifies only the intended motion or interaction, avoiding information leakage and requiring models to maintain physical consistency from the observable evidence in the input scene. We evaluate state-of-the-art video generation models, including both closed-source and open-source ones, using a human-centered protocol based on semantic/action adherence, hidden-relation consistency, and video stability. We define full success as following the requested motion, remaining consistent with observable geometric, optical, and structural constraints, and remaining visually stable. Under this criterion, even the best-performing model achieves only a 65.01% full-success rate. These results show that visually stable videos in this setting often remain incompatible with observable constraints, and VIVG-Bench provides a diagnostic testbed for evaluating and improving reliable video generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.