GRAFT: Transformation-Level Credit Assignment over Prompt Derivation Graphs
Abstract
Despite rapid advances in large language model (LLM) capabilities, effective prompting remains essential for eliciting their full potential, particularly in compound systems governed by multiple interacting instructions. Automatic prompt optimization (APO) automates this process, yet existing methods organize search around prompt candidates and underuse evidence about the transformations that produce them, leading to inefficient exploration under limited budgets. We introduce Graph-based CRedit Assignment For Prompt Transformations (GRAFT), built on the insight that a proposal's evaluation provides evidence not only about the resulting prompt but also about the realized transformation, making the transformation the natural unit of search and credit assignment. GRAFT records optimization as a multi-parent prompt derivation graph and maintains Monte Carlo statistics on transformation edges, allowing accumulated evidence to guide subsequent search toward productive transformations. To capture different notions of transformation utility, we design two reward formulations for GRAFT: mean reward measures average performance, while coverage reward measures how frequently a transformation's output reaches the current per-instance frontier. Across six benchmarks and two open-source LLMs, GRAFT achieves the strongest aggregate performance; on Qwen3-8B, it outperforms GEPA, the strongest baseline, by 3.14 points while using 4.50% fewer rollouts and 20.0% fewer tokens.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.