acceptodds
Under review as a conference paper at ICLR 2027

Reflective Prompt Tuning: From Diagnosing Recurring Failures to History-Aware Revision

Abstract

Large language models (LLMs) have become increasingly capable of following instructions and reasoning through complex problems, making prompting a flexible interface for adapting models without parameter updates. Yet prompt design remains labor-intensive and highly sensitive to formatting, phrasing, and instruction order, motivating automated prompt optimization methods that reduce this manual effort. Existing methods search over prompt candidates or revise prompts from feedback on individual examples or small batches, without maintaining a history of diagnosed failures across prompt versions. As a result, revisions are guided by local feedback rather than by which failures recur across examples and which persist or resolve across iterations. We propose Reflective Prompt Tuning (RPT), a diagnosis-driven prompt optimization framework that grounds each revision in a dataset-level view of the current prompt's recurring failures and in the optimization history. At each iteration, RPT evaluates the target model on the entire optimization set, critiques its incorrect responses, clusters the resulting critiques into recurring failure modes, and returns them with aggregate metrics as a structured report. An LLM optimizer then either revises the prompt, conditioned on this report and a memory of prior reports. RPT also optimizes for confidence calibration, using it in both its diagnostic reports and prompt selection. Across three reasoning tasks and four optimizer LLMs, we find that (i) RPT improves over initial prompts by up to 14.1 points and is the only method competitive on all three tasks, while also improving calibration and keeping prompts compact, (ii) full-set diagnosis, failure clustering, and memory each contribute, and memory lets the optimizer stop once performance plateaus, and (iii) revisions are aligned with the diagnosed failures, though weaker optimizers map distinct failures onto a few generic edits.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.