StructToM-Bench: Locating Where Accuracy Is Lost in Theory-of-Mind Evaluation
Abstract
Theory-of-mind (ToM) benchmarks usually report final-answer accuracy, which can conflate an incorrect answer with a rejected output or, for structured methods, a state that cannot answer the query. We introduce StructToM-Bench, a controlled benchmark of 1000 items with executable gold events, evaluated through direct answering and model-generated event graphs. In a preregistered, blinded two-choice audit, a majority of five human annotators favored the executable gold on each of the 31 items where it disagrees with the source label. For the graph path, strict accuracy separates into three observable stages: output compliance, structured-state availability, and query-result correctness. Across 22 model–deployment configurations of 1–15B parameters, evaluated over five passes against fixed denominators, at least one declared acceptance rule raises the direct score for 10 configurations. Two strict zeros rise to 0.577 and 0.610, while a punctuation-only check leaves 20 of 22 scores unchanged. Among strict graph-path failures, 77.1% are rejected outputs, 17.5% are accepted but unqueryable graphs, and 5.4% are queryable graphs whose computed answer differs from gold. An audit of our structural edge metric found that 30.2% of its gold edges belonged to a family the prompt instructed models to omit, so the metric penalized models for following the prompt. Because positional heuristics selected per operation on this corpus outperform every configuration within each operation, we interpret these findings as measurements of the evaluated pipelines rather than evidence of ToM competence. We recommend reporting strict accuracy alongside acceptance sensitivity, queryable coverage, and scorer and deployment provenance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.