InterMind: What Happens Between Minds? Benchmarking Theory of Mind Across Human–AI Interaction
Abstract
Theory of Mind (ToM), the capacity to attribute mental states to others, supports understanding in human–AI interaction. As an exchange unfolds, each participant interprets and responds to the other, making the relationship between minds central to how the interaction develops. Reasoning within this open-ended process requires models to determine which interpretations and responses the dialogue supports and where that support ends. Existing evaluations leave these complementary judgments and their evidential boundaries insufficiently characterized. We introduce InterMind, a benchmark that examines ToM within the human–AI interaction loop through Mind Reading, Recursive Modeling, Action Forecasting, and Mind Shaping. These tasks connect user-state inference with judgments about the reception of assistant replies, subsequent user actions, and appropriate assistant responses. Questions require models to identify all supported candidates without a disclosed answer count, testing their ability to determine the supported range independently. InterMind contains 700 questions from 568 conversations, grounded in mental-state records and public dialogue and verified through automated assessment and human review. Across 11 models, the best overall exact-set accuracy reaches 56.7%, with differing strengths across tasks. Disclosing the correct answer count raises mean accuracy from 42.4% to 67.6%, revealing a substantial dependence on external constraints when identifying the complete supported set. These findings highlight the challenge of independently forming judgments about an exchange, even when candidate answers are provided.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.