DiPPO: Learning Prompt Proposers with Diffusion Language Model
Abstract
Automatic prompt optimization improves the downstream performance of a frozen target language model without updating its parameters. However, many existing approaches search over prompt strings with a fixed proposer, often relying on strong external language models as proposers or critics, or on human-written initial instructions. We introduce Diffusion Prompt Proposal Optimization (DiPPO), which directly optimizes a masked diffusion language model as a prompt proposer using downstream task rewards. Bidirectional attention in the proposer enables prompt-slot infilling conditioned on the subsequent query. DiPPO samples candidate prompts and updates the proposer using only scalar rewards indicating whether the frozen target model answers correctly. It can discover prompts without a human-written seed instruction or, when one is available, use it as an anchor for preference-based refinement. Across ten benchmarks, the seed-free variant outperforms all methods that do not use a human instruction and exceeds human-written reference prompts on six tasks. The human-anchored variant improves on the human-written instruction on eight tasks, achieves the best performance among all compared methods on four tasks, and records the highest average score, including against methods that use stronger critic models. These results show that a 7B open-weight diffusion language model can not only discover effective prompts from scratch without human-written seeds, but also further improve human-written instructions when they are available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.