acceptodds
Under review as a conference paper at ICLR 2027

Collaborative Theory of Mind: Learning to Self-Revise by Anticipating User Responses

Abstract

Large language models (LLMs) demonstrate strong individual task performance, yet still struggle to collaborate with users over extended, multi-turn interactions, for example by attempting solutions before understanding the user's needs or by failing to revise mistaken assumptions after receiving feedback. We propose Collaborative Theory of Mind (CToM), a framework that addresses these challenges by explicitly training LLMs to anticipate user responses and to use those predictions to improve their own behavior. CToM drafts a response, predicts how the user will respond, and uses this prediction to guide internal revision before presenting a response. During training, the user's reply to each draft serves as the target for predicting the next turn, and we train the assistant to minimize the negative log-likelihood of that reply. We combine this prediction objective with group-relative policy optimization (GRPO) to optimize both task performance and interaction efficiency. Across mathematical reasoning, coding, and writing tasks, CToM achieves the highest collaborative task performance and the best trade-off between task performance and user effort among the evaluated methods. For example, on a collaborative mathematical reasoning task, CToM improves final accuracy from 74.1% to 79.6% compared with naive GRPO, while reducing the number of tokens users need to process by 76%. Behavioral analyses of interaction trajectories show that user-response predictions become more accurate after training, and that more accurate predictions are associated with larger improvements in collaborative task performance through self-revision.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.