MAFIA 1K: A Large-Scale Benchmark for Social Deductive Reasoning from Simulated Game Logs
Abstract
Large language models (LLMs) have demonstrated strong reasoning capabilities across diverse domains, yet their ability to reason about social contexts - where information is asymmetrically distributed and some participants are deliberately deceptive - remains understudied and difficult to evaluate. Social deduction games provide a well-controlled setting for this challenge, but existing benchmarks based on full-length game logs are limited in scale and environmental complexity. We introduce MAFIA 1K, a large-scale social deduction game benchmark comprising 1,000 full-length Mafia game logs across six environments of varying complexity, generated via asynchronous LLM-based social simulations using a Mixture-of-LLMs (MoL) scheme that increases behavioral diversity while maintaining faction balance. The benchmark adopts an observer-style evaluation in which models predict Mafia faction membership from public game logs, isolating social deductive reasoning from confounding factors inherent to interactive gameplay. Our evaluation of 24 LLMs shows that the best model reaches an overall assessment of 43.5%, with top models performing on par with human observers on the same task, and that reasoning models consistently outperform non-reasoning models in every environment. A same-task human evaluation shows that humans achieve meaningful accuracy on both LLM-generated and human-generated logs, with longer response times associated with higher accuracy. An information-source ablation further shows that models rely on both dialogue and voting channels, with different models weighing these channels differently. We release the dataset, evaluation code, and human evaluation data at: https://anonymous.4open.science/r/Mafia-1K-C9B5
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.