acceptodds
Under review as a conference paper at ICLR 2027

Same Message, Same Credit? Diagnosing Credit Stability under Rerouting in Multi-Agent LLMs

Abstract

In multi-agent question answering, determining whether a message's contribution remains valid after communication routes change is essential for reliable system redesign, because outdated scores can misidentify the evidence exchanges on which answers depend. Existing collaboration and attribution methods do not by themselves quantify the error of reusing contribution scores across communication graphs. Alternative routes can make an indispensable message redundant without changing its content. Structural summaries can also merge cases with different effects, preventing predictors from distinguishing their contributions. We propose a three-module diagnostic framework: paired measurement compares retained and disrupted messages on matched tasks, and cross-graph stability computes the optimal fixed credit and its exact worst-case error over evaluated graphs. Exposure sufficiency converts equal-summary effect differences into empirical worst-case error bounds and identifies the responsible states. Measured controlled credits give optimal fixed-score error in all 12 settings, factorial equal-summary bounds are –, and 160 questions from the 2WikiMultiHopQA multi-hop question answering dataset yield a statistically uncertain graph difference in Token-Overlap F-Measure (F1) (95% interval ). These diagnostics distinguish measuring a message's contribution from establishing its reuse, helping developers assess existing scores when communication patterns change.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.