Two-Stage Prompt Optimization: From Reasoning-Guided Search to Gradient-Guided Refinement
Abstract
Existing automatic prompt optimization methods often focus on either broad semantic search or fine-grained local refinement, limiting their ability to benefit from both. We propose a two-stage framework that combines complementary reasoning-based and gradient-based prompt optimization, and introduce GradPO as its second-stage refiner. The first stage can use any general reasoning-based optimizer to make broad prompt improvements based on feedback. In the second stage, GradPO uses target-model gradients to select multiple high-impact spans and combines short, context-aware replacements through beam search, without requiring a separate optimizer model. Experiments on FS-TACRED, FS-FewRel, and MATH-500, spanning relation extraction and mathematical reasoning, show that these two forms of optimization complement each other: local refinement usually further improves prompts found by the first stage. Across the relation extraction tasks, GradPO provides the most consistent improvements, achieving state-of-the-art performance on FS-TACRED with Qwen3-4B while remaining competitive on FS-FewRel. On MATH-500, our GradPO variants improve first-stage prompts more frequently than the alternative refiners, further showing that our approach is effective across both reasoning and non-reasoning tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.