acceptodds
Under review as a conference paper at ICLR 2027

Self-evolving Prompt Optimization for Open-Ended Text with Pairwise Atomic Feedback and Comparative Prompt

Abstract

Prompt optimization offers an effective way to improve large language models (LLMs) for a target task without modifying their parameters, Prompt optimization can improve large language models (LLMs) on a target task without changing their parameters. While most existing methods focus on tasks with well-defined or verifiable evaluation criteria, open-ended text generation poses a harder setting. Responses are often long, lack reliable references, and vary in quality across multiple dimensions, making fine-grained differences difficult for scoring-based LLM judges to detect. Recent methods use pairwise comparisons, but often reduce them to an overall preference or generate feedback only for the losing response. We propose AtomicPO, a self-evolving prompt optimization method that uses pairwise comparisons for both evaluation and feedback generation, without requiring reference texts or sample-specific evaluation criteria. AtomicPO decomposes feedback into atomic strengths and weaknesses across quality dimensions for both responses. During refinement, it also provides a comparative prompt to help preserve useful instructions and guide targeted refinements. Across four open-ended text generation benchmarks, AtomicPO consistently outperforms strong baselines using scoring-based or pairwise LLM judges. These results show the value of atomic feedback from both responses and comparative prompts for optimizing open-ended text generation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.