acceptodds
Under review as a conference paper at ICLR 2027

Beyond Topic Recognition: Social-Function Understanding in Large Language Models

Abstract

Large Language Models (LLMs) have made significant progress in social media mining, yet existing evaluations predominantly focus on surface-level topics and content categories, leaving the communicative functions of online interactions comparatively underexplored. To address this gap, we introduce a theory-informed framework for social-function understanding in social media. Guided by Granovetter’s tie strength theory, Katz’s uses-and-gratifications theory, and Habermas’s public sphere theory, we construct a social-function taxonomy with three top-level categories (Relational, Instrumental, Public) and seven fine-grained subcategories. On this basis, we make three contributions: (1) Reddit-Scene, a multimodal dataset of 6,003 annotated samples from 22 Reddit communities, organized in a three-tier Post–Pillar Comment–Thread (P-R-T) structure with five micro-sociological features; (2) SceneBench, a benchmark evaluating eight frontier LLMs on social-function inference, where the best model reaches only 54.2% seven-class accuracy while a controlled community-identification probe reaches 90.1%, revealing a substantial topic–role gap that is not reliably resolved by additional conversational context; and (3) a controlled multi-agent behavioral probe showing that social-function information has an observable influence on downstream agent responses. A fine-tuned encoder reaches 76.2%, and a three-annotator human baseline reaches 78.3% on social-function inference but 84.7% on community identification, further suggesting that the main difficulty lies in social-role inference rather than surface community recognition.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.