Select or Create? Behavioral Preferences and Quality Gains in LLM Humor Tasks
Abstract
Do large language models (LLMs) prefer to select existing content or create new alternatives, and are such behavioral preferences associated with quality gains? We investigate selection-creation preferences in a news-based humor task, in which models evaluate candidate jokes and decide whether to adopt one unchanged or generate a new alternative. We introduce JEST, a multilingual dataset comprising 3,923 news headlines in English, Spanish, and Chinese, each paired with four jokes generated independently by distinct LLMs. We further propose an evaluation framework that conditions quality comparisons on models’ initial decisions. Across 11 state-of-the-art LLMs, we uncover distinct behavioral preferences, with average selection rates ranging from 33.7% to 99.8% and model rankings proving highly consistent across language datasets. Subsequent quality evaluation reveals that frequent selection does not guarantee accurate identification of the highest-scoring candidate. Forcing selection-preferring models to create does not consistently yield jokes that outscore their initial selections. For GPT-5.4, forced generations score higher than its selections but still lower than the highest-scoring candidates on average. Similarly, although some creation-preferring models generate jokes that score above the candidate mean, none achieves an average score as high as that of the highest-scoring candidates in any language. Overall, our findings highlight the need to jointly evaluate models' abilities to generate new content, identify high-quality candidates, and decide when to create.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.