acceptodds
Under review as a conference paper at ICLR 2027

BehEvolve: Behavior-Guided Evolution of Multi-Agent Program Policies under Sparse Task Feedback

Abstract

Large language models (LLMs) can synthesize executable multi-agent policies, but improving these policies through program evolution requires informative feedback that distinguishes promising candidates. Under sparse task-completion feedback, unsuccessful programs receive identical fitness scores, creating zero-fitness plateaus where fitness alone cannot determine which programs to expand or how they should be revised. We present BehEvolve, an idea–program evolution framework that leverages execution behavior to guide program evolution without requiring hand-designed process rewards. The framework retains parent–child code changes and paired execution traces, from which a feedback model produces behavioral diagnoses for program mutation and relative preferences for parent selection. Behavioral progress further guides idea-level scheduling when a strategy branch stalls. Across eight VMAS tasks and five Overcooked layouts, BehEvolve achieves the strongest average performance, reaching an average success rate of 81.5% on VMAS and 12.6 soups on Overcooked. It outperforms the strongest general evolution baseline, AdaEvolve, and the strongest Overcooked-specific baseline, KnowPC, by 29.0% and 18.9%, respectively. Further analyses show that the three behavioral feedback pathways—program mutation, parent selection, and idea-level scheduling—work together to improve search efficiency and final performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.