acceptodds
Under review as a conference paper at ICLR 2027

Appearance-free Video Generation for Measuring a Perception Gap in Motion-Defined Forms

Abstract

VLMs have achieved remarkable progress on video understanding benchmarks, yet it remains unclear whether they are truly as capable as humans. A key aspect of human perception is persistence of vision. We can identify motion patterns in videos whose individual frames appear as pure noise (without any appearance or recognizable content). To quantify this capability for VLMs, we propose an algorithm to generate appearance-free videos from any given source of motion. Each frame is indistinguishable from noise, and only motion contains the information. We then introduce AppearFreeVQA, a benchmark with ≈5,300 appearance-free videos spanning four motion categories and three difficulty levels. We demonstrate that humans achieve approximately 93% accuracy, while frontier VLMs (GPT-5.1) reach only 41.7%, slightly above random chance. Oracle experiments show that supplying target optical flow improves Qwen3.5-9B from 14.6% to 66.8%, indicating that motion extraction is a major bottleneck. Finally, we show that training on AppearFreeVQA can improve VLMs' capabilities on camouflaged object detection, a real-world application where appearance cues are limited.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.