acceptodds
Under review as a conference paper at ICLR 2027

HiP-Bench: A Benchmark for Hidden-Picture Optical Illusion Generation

Abstract

Visual generative models have demonstrated impressive capabilities in synthesizing high-quality images from multimodal inputs. However, their ability to generate hidden-picture optical illusions, where hidden targets are integrated into background images and emerge only under specific viewing conditions, remains underexplored. To fill this gap, we propose a novel benchmark for Hidden-Picture Optical Illusion Generation (HiP-Bench). Specifically, we construct a dataset of 512 samples and use 35 models to produce such illusions across three paradigms: (i) text-to-image, (ii) instruction-based image editing, and (iii) image-conditioned generation. A panel of human experts then evaluates 13,824 generated images across four dimensions, primarily focusing on illusion fidelity and shape consistency. Through five research questions, our analysis reveals the capability boundaries of existing models. Moreover, we explore Multi-Scale Distance Simulation (MDS), an automated baseline based on multimodal large language models that assesses illusion effectiveness through multi-scale visual inputs. Overall, this study establishes a reliable benchmark to facilitate and evaluate future progress.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.