Discovering Multiagent Learning Algorithms with Large Language Models
Abstract
Much of the advancement in Multi-Agent Reinforcement Learning (MARL) for imperfect-information games has historically depended on the manual, iterative refinement of algorithmic baselines. Recently, evolutionary coding agents powered by Large Language Models (LLMs) have emerged as powerful tools to automate this discovery process. In this work, we deploy one of such agentic frameworks, AlphaEvolve, to navigate the design spaces of two distinct game-theoretic paradigms: counterfactual regret minimization (CFR) and policy-space response oracles (PSRO). This automated search yielded two algorithms: Volatility-Adaptive Discounted (VAD-) CFR and Smoothed Hybrid Optimistic Regret (SHOR-) PSRO, which are consistently competitive with state-of-the-art human-designed baselines across an 18-game evaluation suite spanning Poker, Goofspiel, Liar's Dice, Blotto, and Battleship variants. However, because the LLM optimizes for fitness on a specific training set, it often constructs highly synergistic, complex mechanisms tailored to those environments. Through systematic train-only ablation criteria, we distill these discoveries into two minimal closed-form solvers: Warm-started Optimistic Predictive (WOP-)CFR and Projection Matching (PM-)PSRO. Ten independent search runs confirm that the distilled cores recur while the overfitted machinery differs across runs. Distillation isolates this stable algorithmic core, achieving superior held-out performance at drastically reduced structural complexity: WOP-CFR improves on scaled variants within family and maintains parity out of family, while PM-PSRO improves in both regimes and most strongly out of family.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.