acceptodds
Under review as a conference paper at ICLR 2027

Sharp Norm Conditions for Logit Fixed Points

Abstract

Entropy regularization is often used in multi-agent learning to reduce instability caused by agents influencing one another. A natural intuition is that learning should be easier to stabilize when a change in one agent's behavior has only a limited effect on the others. In this paper, we show that the size of a payoff change alone can be misleading: even a large change may leave an agent's choices unchanged. What matters is whether the change alters the relative payoffs of different actions. Building on this observation, we derive a new sufficient condition for stability of logit responses that accounts only for changes that affect choice, and specifies how much regularization is enough. Under this condition, responses converge to the same fixed point from every initial state. A two-action example establishes the sharpness of the condition's universal constant: immediately beyond the bound, the system can have several distinct long-term outcomes. Experiments show that bounds based on the full payoff matrix can count changes that leave choices unchanged, requiring unnecessarily strong regularization or failing to guarantee convergence in systems that do converge. The new condition avoids this problem. We believe these findings help explain stability in multi-agent learning and guide the choice of regularization strength.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.