acceptodds
Under review as a conference paper at ICLR 2027

Human-Level on Which Questions? A Ranking Reversal in Visual Reasoning

Abstract

As frontier VLMs increasingly achieve near- or super-human performance on specific benchmarks, there has been increasing speculation that VLMs have superhuman capabilities, based on the assumption that performance should generalize. We think that it is only healthy to question this assumption. We find that 8 frontier VLMs match or exceed human performance on AI-generated visual-reasoning questions (AIGEN) but fall substantially behind humans on human-generated questions (HUMANGEN). This happens because frontier VLMs are sensitive to the difference between AIGEN and HUMANGEN, while humans are robust to it. We carry out an in-depth analysis of both audits to see what causes this trend. We find that surface-level phrasing does not explain it. Instead, the result is due to differences in the reasoning paths required by the two audits. HUMANGEN tends to have longer paths that involve finding visual cues that frontier VLMs are less likely to look for. Providing these cues in oracle form restores performance for frontier VLMs, but to varying degrees, suggesting that some VLMs still lack world knowledge necessary to deduce an answer from these cues.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.