PolyCross: A Benchmark for Identifying Hidden Meanings in Short Phrases with LLMs
Abstract
Identifying hidden meanings in short text snippets is a challenging task with applications in legal documents, safety-critical technical documentation, creative writing, and word puzzles. In this work we investigate the capacity of large language models (LLMs) to identify concealed meanings in short phrases, provide rationales for the existence of meanings identified, and translate phrases with multiple meanings into another language. For the empirical evaluation we introduce POLYCROSS, a benchmark dataset built from a special class of crosswords clues in Romanian which we call polysemantic clues. Each such clue carries multiple meanings, including a subtle meaning that leads to the actual crosswords solution. The baseline is THEMCROSS, a dataset of general-knowledge Romanian crosswords clues with no hidden meanings. We use the two data sets to evaluate three LLMs (GPT-OSS, LLaMA and Granite) with progressively stronger hints about solutions. We assess LLM accuracy of finding crosswords solutions and the quality of LLM rationales to explain their solutions. Our results show that the performance of modern LLMs is substantially lower on polysemantic clues (POLYCROSS) than on the baseline set (THEMCROSS). This highlights an important gap in current LLM capabilities and establishes POLYCROSS as a benchmark to measure future progress on this problem.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.