acceptodds
Under review as a conference paper at ICLR 2027

MedDAC: Learning Directional Communication via Correction and Harm Modeling for Medical Multi-Agent Reasoning

Abstract

Multi-agent reasoning relies on communication among agents, yet peer messages can either correct an erroneous receiver or mislead an initially correct one. Moreover, agents may revise their answers through self-reflection alone, making it difficult to distinguish peer influence from ordinary second-pass reasoning. We present MedDAC, a framework for learning directional agent communication policies by modeling both corrective and harmful peer effects. MedDAC evaluates receiver-specific sender-subset interventions against matched self-review to estimate the incremental local and system-level effects of communication actions. MedDAC-Static aggregates these effects to learn a fixed sparse communication graph under a message budget, while MedDAC-Dynamic predicts question-conditioned communication actions from the current multi-agent state. Across five medical QA benchmarks, the complete MedDAC-Static pipeline achieves 74.57% macro-average accuracy, compared with 69.77% for single-agent inference. Under the same Majority Vote aggregation, its six-message policy achieves 73.13%, comparable to the 72.99% obtained with 12-message Full Communication. These results show that modeling both corrective and harmful peer effects enables selective communication while maintaining competitive reasoning accuracy with substantially fewer directed messages.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.