acceptodds
Under review as a conference paper at ICLR 2027

From MARL Policies to Swarm Laws: The Compressibility Spectrum of Learned Coordination

Abstract

Swarm intelligence and multi-agent reinforcement learning (MARL) are two com- mon paradigms for multi-agent coordination, yet they construct collective behavior from opposite directions. Swarm rules such as Boids and mean-shift consensus are local, interpretable, and independent of population size in their formulation. However, they must be designed by humans. MARL policies instead learn co- ordination from environmental rewards and can adapt to different tasks through training. However, the learned policies remain black boxes and may degrade when deployed beyond their training population. These complementary strengths mo- tivate a more specific question. When can a trained collective policy be distilled into a compact local controller that preserves useful organization and transfers across population sizes? Our results show that the answer depends jointly on the policy, its observables, and the rule language. A causal task makes the information boundary explicit: when a hidden bit determines the correct target among locally indistinguishable options, blind mining falls to coin-flip performance, while expos- ing that bit restores the teacher’s performance curve. In SMACv2, three unrelated compressors reveal a sharp Protoss–Terran contrast. On the compressible side, mined local laws preserve learned organization and transfer beyond the training population. In continuous pursuit–encirclement, a sixteen-number law exceeds its own teacher and maintains capture performance up to 32× the training population at constant density. We further remove the hand-designed vocabulary and recover compact executable programs from a frozen grammar of 208 generic local terms. The recovered structure changes systematically across pursuit, coverage, and public MPE Predator–Prey tasks. Finally, direct return optimization and mining from dif- ferent teachers separate rule-class capacity from teacher-specific content. Together, these results characterize when learned collective behavior can be compressed into a local swarm law, connecting MARL’s automatic discovery of coordination with the interpretability and population scalability of swarm intelligence.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.