Beneath the Surface: Investigating LLMs' Capabilities for Communicating with Subtext
Abstract
Human communication is fundamentally creative, and often makes use of subtext—implied meaning that goes beyond the literal content of the text. Here, we systematically study whether language models can use subtext in communicative settings, and introduce three new evaluation suites to assess these capabilities. Our evaluation settings range from writing & interpreting allegories to playing multi-agent and multi-modal games inspired by the rules of board games like Dixit. We find that out of the box, frontier models exhibit a bias towards overly literal, explicit communication. However, a shared common ground between agents can significantly help reasoning models like Gemini-2.5-pro and GPT-5 to communicate subtext to an intended audience sharing the common ground. In our environment Visual Allusions, we find 25%-40% (absolute) reduction in literal clues with a common ground of short stories. Similarly, in our storytelling environment The Aesopian Author, an author agent's success rate in communicating a forbidden idea to a critic, which an inquisitor interprets as benign, drops from 54% to 28% when critic loses the common ground with the author. In The Aesopian Author, we also find stronger models can convey subtext that a weaker monitor model might miss. For allegory understanding, we find paratextual and persona conditions to significantly shift the interpretation of subtext. Overall, our work provides quantifiable measures for an inherently subjective phenomenon like subtext and reveals both strengths and weaknesses of current LLMs in these settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.