RBEPO: From Root-Cause Diagnosis to Branching Evolutionary Prompt Optimization
Abstract
With the widespread deployment of large language models in agentic systems and domain-specific tasks, using execution feedback to enable system self-evolution has attracted increasing attention. Automatic prompt optimization provides a lightweight path for systems to continually improve task performance without updating model parameters. However, identifying problems from execution feedback does not necessarily reveal effective repair strategies; joint rewrites across multiple directions obscure the actual effects and reuse potential of individual repairs. We propose RBEPO (Root-cause-guided Branching Evolutionary Prompt Optimization), a prompt optimization framework centered on cross-instance failure modes that jointly organizes targeted branching search and mechanism-level experience accumulation. RBEPO infers repair directions through per-instance error analysis and cross-instance root-cause clustering, generates and evaluates targeted rewrite branches for individual failure modes, and integrates effective, complementary repairs. Building on these evaluated rewrites, the method distills rewrite attempts and their evaluation results into cross-iteration experience to guide subsequent rewrites. To control the resulting candidate-evaluation cost, prompt-edit fragility-guided racing is introduced. Across multiple tasks and models, RBEPO consistently achieves higher test scores than the compared baselines: its average score is approximately 3.35 points higher than that of the strongest baseline in aggregate, while its final prompts are on average only about 1/6 as long. Comparisons of evaluation strategies further show that this strategy reduces the total sample-evaluation volume to about 38% of full evaluation while maintaining performance close to that of full evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.