Psy-MARL: Psychology-Guided Multi-Agent Reinforcement Learning For Psychiatric Interviewing
Abstract
Psychiatric interviewing plays a central role in diagnosis and treatment, requiring clinicians to explore symptoms while attending to the patient's psychological state and communication needs. Advances in large language models offer new opportunities for psychiatric interviewing, yet clinical dialogue reasoning in this highly context-dependent setting remains challenging. To this end, we propose Psy-MARL, a psychology-guided multi-agent reinforcement learning framework inspired by Theory of Mind and Motivational Interviewing. Psy-MARL comprises three agents: the Belief agent assesses the patient’s state, the Strategy agent plans a communication strategy, and the Response agent generates a physician response. Their policies are jointly optimized through reinforcement learning. To obtain informative reward signals and distinguish the contributions of agents, we introduce preference-based coalition credit assignment. This method models role collaboration as a cooperative game, fits coalition utilities from LLM preferences using a Bradley-Terry model, and derives role-specific rewards through Shapley allocation. We train and evaluate Psy-MARL on real-world psychiatric consultation dialogues, using LLM-based pairwise evaluations to assess alignment with recorded physician responses and preference under clinical evaluation criteria. Psy-MARL achieves positive net-win margins against all evaluated single-agent and multi-agent baselines under both protocols, including clinical margins of 6.36-27.19 percentage points against three multi-agent reinforcement learning methods. Ablation studies support the benefits of preference-based coalition utility estimation and role-specific credit allocation. These findings support psychology-guided agent collaboration and preference-based joint training as a promising approach to strengthening clinical reasoning in psychiatric interviewing.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.