Honkaku-Bench: Can AI solve fair-play murder mysteries?
Abstract
Evaluating whether a large language model can reason, rather than rely solely on retrieval or pattern matching, requires assessing not only its final answer but also the validity of the reasoning steps supporting it. Honkaku mysteries offer a natural setting for this evaluation as fair-play puzzles challenge readers to explain seemingly impossible events using clues in the story. We introduce Honkaku- Bench, a benchmark of 70 mystery puzzles for evaluating evidence-grounded reconstruction, paired with unique official solutions written by human experts. Given a complete case file and any accompanying diagrams, models must produce an open-ended solution explaining what happened, how it was possible, and which evidence supports the reconstruction. Across eight reasoning models evaluated through a common harness, GPT-5.6-sol and Fable 5 recover 61.3% and 59.3% of rubric credit on average but receive full credit on only 6 and 8 of 70 cases. All eight models score 18.1–26.8 points lower on the cases with visual evidence than on the text-only cases. Our case studies reveal how partial progress fails to build a coherent reconstruction: a model may recover the mechanism while naming the wrong culprit, or propagate one misread clue through consistent arguments. Honkaku-Bench thus tests whether models can turn scattered textual and visual clues into one explanation consistent with the entire case.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.