acceptodds
Under review as a conference paper at ICLR 2027

DO VISION-LANGUAGE MODELS SEE ILLUSIONS LIKE HUMANS DO? A PARAMETRIC PSYCHOPHYSICS BENCHMARK FOR VLM VISUAL BIAS

Abstract

Vision-language models (VLMs) are routinely evaluated on optical illusions by scoring a single modified image “right” or “wrong,” collapsing a graded perceptual phenomenon into a binary label. We introduce ILLUSIONCURVE-BENCH, a benchmark that instead sweeps four classical illusions (Ebbinghaus, M¨uller-Lyer, Ponzo, Vertical–Horizontal) across 11 stimulus-magnitude steps each, fits psychometric functions to three open-weight VLMs’ (Qwen2.5-VL-3B/7B-Instruct, LLaVA-1.5-7B) forced-choice responses, and compares the resulting point-ofsubjective- equality (PSE) and slope estimates to published human norms – with the A/B label and position assignment counterbalanced across repeats so a model’s response-format bias cannot masquerade as a genuine perceptual offset. Of the 12 model–illusion response curves, only three are classified graded/human-like (one on a sparse cell where 86% of responses were “Equal”), while most (nine) are erratic or fail to converge at all, and 4 of the 8 that do converge place their PSE outside the tested range entirely. Doubling our repeat count mid-study to rule out sampling noise made this picture more, not less, extreme, and an auxiliary same/different confidence probe is similarly decoupled from stimulus ambiguity and unstable across reruns – one 7B-scale model stays within 0.25 of a constant “Same” verdict on every illusion, and which cells look dip-like at the point of maximal ambiguity changes across every rerun. Our results indicate that binary-accuracy illusion benchmarks systematically misattribute categorical, memorization-driven behavior to human-like perceptual bias, and that psychophysical curve-fitting with response-format counterbalancing is necessary – though not sufficient on its own, given the resampling instability we observe – to tell the two apart. We release ILLUSIONCURVE-BENCH, our sweep generators, and evaluation code.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.