What Should the Reflector See? An Empirical Study of Evidence in Reflective Prompt Optimization
Abstract
Reflective prompt optimization revises instructions using examples of a model’s behavior, but which evidence the reflector should receive remains unclear. We compare nine reflection strategies within a single-parent Pareto-guided search, varying evidence composition, visibility of examples, candidate selection and domain-knowledge policy. Using Qwen3.5-9B as both task model and reflector, we complete 45 optimization runs on five datasets and report test accuracy averaged over three final evaluations. No strategy dominates across tasks. Failure-only evidence produces the largest average improvement, driven by a 33.7-point gain on LiveBench-Math, whereas balanced evidence and validated blind generation share the best mean accuracy rank. Balanced evidence is not uniformly better than failure-only evidence, and explicit knowledge retention helps multi-hop question answering but hurts some other tasks in these runs. Empirical selection with natural evidence outperforms an LLM judge on average, although their computational costs differ. The results highlight task dependence and the need to distinguish calibration improvement from held-out performance, with repeated optimization runs still needed for stronger conclusions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.