Words at Play: Benchmarking Audio Pun Understanding in Large Audio-Language Models
Abstract
Puns represent a typical linguistic phenomenon that exploits polysemy and phonetic ambiguity to generate humour, posing unique challenges to the understanding of natural languages. Compared with text and images, spoken language plays a more central role in human communication. However, datasets and systematic resources for audio puns remain scarce, leaving this important modality largely underexplored. In this paper, we present APUN-Bench, the first benchmark dedicated to evaluating large audio language models (LALMs) on audio pun understanding. Our benchmark contains 4,447 audio samples annotated across three stages: pun recognition, pun word location and pun meaning inference. We conduct a deep analysis of APUN-Bench by systematically evaluating 10 state-of-the-art LALMs, uncovering substantial performance gaps in recognizing, localizing, and interpreting audio puns. Motivated by our analyses of positional biases in pun localization and error patterns in pun meaning inference, we further propose a prompt-based strategy and a fine-tuning framework to improve audio pun understanding. Our findings provide actionable insights toward advancing humour-aware audio intelligence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.