acceptodds
Under review as a conference paper at ICLR 2027

Restoring Co-Adaptation: An Endpoint Audit of Asymmetric Training in Multi-Agent Market Making

Abstract

When an endpoint produced by asymmetric multi-agent training is evaluated with all policies frozen, the evaluation disables the learning responses whose suppression helped shape that endpoint. We distinguish frozen persistence from algorithm- and horizon-relative co-adaptive response stability and introduce a restoration audit: resume all suppressed update channels from the learned endpoint and compare the resulting system-statistic change with a same-budget continuation from a co-adaptive reference. In a controlled three-agent market-making testbed, staged train-and-freeze training increases mean minimum quoted spread (MQS) relative to simultaneous co-adaptation by (95% CI ; 20/20 seeds positive). After 600 restored episodes, the mean recovery fraction is (95% CI ), whereas the matched simultaneous continuation drifts by only (95% CI ). A factorial audit shows that the tested spread shift is associated with the interaction of long active-agent blocks and elevated placeholders (primary interaction , 95% CI ; 20/20 seeds positive). The disadvantage follows the first-trained role across identities, and Independent DQN learners reproduce the prespecified block, placeholder, and restoration contrasts in the same testbed. Restoration is not an equilibrium or exploitability test; within this controlled setting, it shows that frozen persistence can overstate stability under the specified resumed-learning dynamics.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.