acceptodds
Under review as a conference paper at ICLR 2027

Moral Hazard in Multi-Agent Language Models

Abstract

Cooperation can fail when costly, hard-to-observe effort benefits another agent. The Dialogue Moral Hazard Game instantiates this incentive problem for language agents: an agent chooses between a local reward and paying to reveal a hidden safety fact for a successor's decision. Across fourteen open-weight and four frontier models, we distinguish information acquisition, communication, downstream use, and team success. Private-share interventions reveal both accurate incentive-boundary tracking and safety-prioritizing departures despite correct payoff comprehension. Outcome-directed prompt optimization can bypass the cooperative mechanism: GEPA raises Muse Spark 1.1's team success from 22.2% to 100% while reducing querying from 51.1% to 0.3%. Freezing its prompts and changing the rank–label mapping reduces success to 12.5% and then 0%. We introduce CREDIT (Counterfactual Replay for Evidence-Driven Information Transfer), a mechanism-aligned multi-agent prompt-optimization algorithm using matched hidden-state twins, complete action replay, and staged causal credit. Across five models with three optimization seeds each, CREDIT increases realized information transfer relative to standard GEPA. Muse and Claude Opus sustain perfect query-mediated success across mappings; open-weight models expose remaining acquisition and downstream-use bottlenecks. These results distinguish successful outcomes from the information mechanisms that produce them.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.