Which Messages,at What Budget? A Complete-Subset Audit of Learned Message Sharing
Abstract
In multi-agent scientific reasoning, a learned selector decides which of several proposed solutions reach the model that writes the final answer. When that choice leaves accuracy unmoved, accuracy alone cannot tell you why: the developer cannot see whether a winning combination was unavailable, available but not selected, or selected but not repeatable. We separate the three by recording what the answering model does under all 32 subsets of five fixed messages, across 975 scientific questions. The result is a diagnosis rather than a verdict. Winning combinations turn out to be plentiful, as some subset succeeds on 91.2% of biology questions against 70.9% for answering alone, yet neither of two independently seeded selectors converts that headroom into a significant gain on any of five cohorts. Within the selector family we evaluate, the shortfall is in selection rather than in what was available to select. The same tables show that a harmful message rarely forces a rebuild: 245 of 255 affected questions still admit an ordering that is optimal at every budget. But structure is easy to overstate, and our own controls temper it: four of five cohorts match a density-preserving null, and the one exception does not recur on fresh executions. This study therefore suggests a practical order of work: diagnose which component limits the pipeline before increasing its message budget. Absent combinations implicate the candidate generators or the answering model; available but unselected ones implicate the selection policy; unstable ones implicate re-execution variance and receiver alignment. Evaluation should indicate which component to improve next, rather than motivate more messages.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.