acceptodds
Under review as a conference paper at ICLR 2027

MUNCH: Benchmarking Multimodal Humor Understanding in Visual Language Models through Meme Comprehension

Abstract

A meme template is a reusable image that stands for a situation or stance, and a caption turns it into a joke by filling that slot with a new case. Recognizing which caption fills a template is an abstract form of multimodal humor understanding that evaluations of vision–language models rarely isolate. We introduce MUNCH, a benchmark of 123K Reddit memes whose captions have been inpainted out and recast as four-way multiple-choice questions, with test partitions that hold out templates and entire semantic clusters. Across 23 zero-shot multimodal large language models (MLLMs), the strongest scores 14 to 21 points below our human performance estimate, although these models describe images fluently. A contrastive model fine-tuned on MUNCH reaches the human performance estimate on templates seen in training (86%), but its gains shrink as templates and their meanings move away from training, and on held-out semantic clusters it scores 13 points below the estimate (62%). Our results show that multimodal image–text humor understanding is challenging for otherwise highly-performant models but can be learned with training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.