AutoPEER: Issue-Guided Prompt Evolution for Long-Horizon Agents
Abstract
Large language models (LLMs) increasingly power agents that solve complex tasks through multi-step tool interactions. Existing prompt optimization frameworks learn from execution feedback, but give limited attention to credit assignment in long-horizon tasks. Effective prompt evolution requires identifying which decisions in a trajectory should be improved and updating the prompt accordingly. We introduce AutoPEER, a framework that combines Issue-level credit assignment with progressive behavioral verification. AutoPEER locates decisions that need better guidance and records them as Issues to guide targeted prompt patches. Progressive verification tests whether patches change the intended actions, improve complete-task performance, and generalize across tasks after merging. A multidimensional reward rubric selects prompts based on task success and execution quality. Experiments on AppWorld show that AutoPEER improves both task completion and interaction efficiency across models, achieving the highest completion on each evaluated split. These improvements are observed with and without training reference solutions. Notably, prompts learned with one model improve another without further optimization. Additional evaluation on the telecom domain of τ²-Bench demonstrates gains in a different tool environment. Ablations show that targeted repair and progressive verification help agents solve more tasks, while multidimensional rewards encourage more efficient execution.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.