Structura: Intrinsic Scoring and Inference-Time Refinement against Structural Hallucinations in Video Generation
Abstract
Modern video generators produce visually compelling clips, yet objects can develop impossible structures as they move: limbs fuse or multiply, parts detach, and geometry changes inconsistently across frames. Mitigating these spatiotemporal structural hallucinations requires feedback that identifies emerging failures and guides correction during sampling. We introduce Structura, a training-free framework that obtains this feedback from the generator’s own reconstruction dynamics. We slightly re-noise each intermediate clean prediction and reconstruct it, using low-noise reconstruction as a local geometric probe. Unreliable regions tend to change more under this probe, and large changes recur at the same locations when a structural failure persists across sampling steps. Structura accumulates the magnitude of these changes to stabilize spatiotemporal localization, while retaining their signed displacement to guide stochastic local resampling. We introduce Structura Bench, a collection of hallucination-prone cases for evaluating video structure. Experiments on Dynamic-Bench and Structura Bench show improved structural and anatomical correctness across video generators. Applications to video object removal and world action models further demonstrate the use of reconstruction feedback for temporal refinement and action completion.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.