ThreadBench: Benchmarking Realism in Agentic Simulation of Real-world Discussions
Abstract
LLM agents are increasingly used to simulate real-world interactions, but it remains unclear whether simulated behaviors preserve the content patterns and interaction dynamics of real human behaviors. Existing evaluations often focus on individual comments or task performance, rather than whether simulated behaviors match human behavior at thread level. In this paper, we use Reddit discussions as a concrete testbed for real-world social simulation. Reddit provides public, topic-grounded, multi-party interactions with diverse opinions, experiences, and interaction patterns, making it a useful setting for evaluating whether LLM agents reproduce the distributional and structural properties of real online discussions. We introduce ThreadBench, a benchmark for Reddit discussion simulation built from 9,817 real Reddit threads spanning 8 diverse domains and 116 subdomains. ThreadBench uses statistical analyses to compare generated and real discussions across 12 metrics grouped in 4 families: uniformity, expression, tone, and interaction. We further develop Thread-Prompt-based Discussion Generation (ThreadPDG), a Planner–Writer framework that learns how a real discussion grows. The Planner produces thread- and comment-level plans from a real seed discussion, and the Writer generates comments from these plans. For each seed, we generate multiple candidate threads and combine their best branches using thread-level statistics. Across all domains and multiple LLMs, ThreadPDG achieves ThreadBench pass rates of 58–100%, compared with 0–17% for the baselines. Human evaluation with 436 annotators further shows that ThreadPDG is not distinguished from real discussions in 78.7% of pairs, outperforming existing simulation methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.