acceptodds
Under review as a conference paper at ICLR 2027

SpatialSherlock: Diagnosing Spatial Inconsistencies in AI-Generated Videos

Abstract

Diagnosing spatial inconsistencies in AI-generated videos requires distinguishing changes explained by camera motion, object motion, and occlusion from conflicts with the depicted scene evolution. We introduce a data pipeline that connects spatial skills reconstructed from public datasets with structured human diagnoses of generated videos. The annotations specify affected regions, required spatial corrections, and explanations, making the spatial basis of each judgment separately assessable. These references support both diagnostic reasoning supervision and component-level reinforcement-learning feedback. The pipeline curates 92,000 examples across four diagnostic tasks from more than 20,000 generated videos. Using this supervision, we train spatialsherlock, based on Qwen3.5-9B, through supervised fine-tuning and joint reinforcement learning. The model forms local diagnoses and checks them against available intervening frames before aggregating video-level judgments and temporal evidence. spatialsherlock achieves 81.3% accuracy on frame-pair diagnosis and 78.7% on source-video diagnosis, exceeding the strongest tested commercial baselines by 40.3 and 16.7 percentage points, respectively. It also improves on its backbone across all nine external spatial benchmarks evaluated zero-shot, with gains of 1.6–16.6 percentage points.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.