acceptodds
Under review as a conference paper at ICLR 2027

SAGE: Self-Attribution-Guided Evolving

Abstract

Prompt evolution optimizes the instruction given to a frozen language model: it maintains a population of candidate instructions and repeatedly mutates, evaluates, and selects them. Existing prompt optimizers select candidates by task accuracy alone, but accuracy cannot tell whether a correct answer is due to the instruction. Two instructions can be equally accurate even when one induces a distinct solution strategy and the other merely rewords an existing candidate and leads the model to behave the same way. We formalize this blind spot: among candidates whose rewards follow the same distribution, the rewards carry no information about which candidate produced them, whereas the model's trajectories do. To recover this information, we introduce self-attribution, which measures how much better a candidate's own instruction explains its trajectory than the best competing instruction does, and we prove that it yields a computable lower bound on the information that accuracy discards. We then propose SAGE (Self-Attribution-Guided Evolving), which adds this signal to two decisions of GEPA, a state-of-the-art prompt optimizer: which candidates are selected as parents, and which mutated children are kept. Across four open-weight models spanning a range in size and six tasks, SAGE achieves the highest average accuracy on every model and improves over GEPA by up to 4.0 points. It also ends below its starting prompt in only 3 of 24 model–task pairs, compared with 12 for GEPA. These gains persist on harder out-of-distribution problems, under cross-model prompt transfer, and with a black-box API model whose self-attribution is computed by an open-weight surrogate.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.