Do Generative Priors Align with Human Naturalness Perception?
Abstract
Visual generative models learn probability distributions over natural images, but whether their native priors directly capture human perception of image naturalness remains unresolved. Here, we probe these priors through native prediction errors across 25 open image and video generators. Because raw single-image losses are dominated by scene content and visual complexity, we evaluate directional loss differences using content-preserving, paired relational interventions that selectively disrupt facial configurations or physical illumination consistency while limiting changes in low-level image statistics. In representative conditions, these loss differences reproduce selective human sensitivities and tolerances, capturing the classic Thatcher effect on faces and shape-dependent responses to illumination inconsistencies. Crucially, they reliably track continuous human naturalness judgments across individual stimulus pairs (peaking at on faces and on physical scenes). Across the evaluated models, overall sensitivity to these violations broadly covaries with human alignment, yet the two dissociate along denoising schedules, with fine-grained perceptual alignment typically peaking before sensitivity. Importantly, generative loss differences retain positive partial correlations with human judgments after controlling for feature distances from frozen vision encoders and standard image quality metrics. Together, these findings demonstrate that learning visual distributions yields generative loss landscapes that capture distinct aspects of human naturalness perception.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.