acceptodds
Under review as a conference paper at ICLR 2027

SKARL: Kernel Mean-Field Reinforcement Learning for Zero-Shot Population-Size Transfer

Abstract

Large-scale multi-agent reinforcement learning (MARL) requires policies that scale to many agents and tolerate population mismatch between training and deployment. Mean-field reinforcement learning (MFRL) provides a size-independent interface, but standard moment summaries discard distributional structure. We introduce **S**calable **K**ernel Me**A**n-Field Multi-Agent **R**einforcement **L**earning (SKARL), which embeds empirical populations by kernel mean embeddings and represents individual values by cylindrical kernel functionals. Under a characteristic kernel, we prove density of this function class. Under same-law sampling, we further bound the distance between the laws of empirical embeddings at sizes and by a term proportional to , yielding the same rate for a fixed Lipschitz cylindrical functional. This result isolates representation sensitivity rather than end-to-end return preservation. For smooth kernels, we derive an RKHS Riesz gradient and its projection onto a fixed-size Nystr\"om dictionary. Across variable-size MPE, MAgent2, and PettingZoo tasks, SKARL scales to large populations and transfers zero-shot. In MAgent2 pursuit, training with and testing with up to improves over first-moment MFRL by - return per agent; in combined arms, training with retains positive mean return from to , while all evaluated baselines are negative.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.