RB-Prompt: Detecting and Repairing Prompt Saturation in Reflective Optimization
Abstract
Prompt optimization adapts large language models to new tasks without changing their weights. Reflective optimizers rewrite prompts using feedback on model errors, but can keep adding instructions after performance stops improving. We call this prompt saturation: search spends evaluation budget on longer prompts with little gain. We address this problem with RB-PROMPT, which detects saturation from recent scores and instruction growth. It returns to a branch’s best saved prompt and asks reflection to remove or combine unnecessary rules. Search continues with a penalty for added instruction length and Bayesian selection guided by past rewrite gains. We evaluate on HotpotQA, GSM8K, and an existing document extraction task at a major global automotive manufacturer, using three model families. Under matched task-evaluation budgets, RB-PROMPT achieves the highest or joint-highest score in all nine settings, with the largest gains on HotpotQA and extraction. It achieves higher scores while using 40–45% fewer optimization tokens than TextBO-GEPA. Comparisons from the same saved states isolate the repair’s benefit over continued reflection. We show that adding repair to GEPA and TextBO-GEPA improves performance without changing their parent-selection rules. These findings support adaptive branch repair as a reusable component of reflective prompt optimization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.