acceptodds
Under review as a conference paper at ICLR 2027

Differential Prompt Optimization: Pinpointing Student Errors via Teacher Audits

Abstract

Automatic prompt optimization typically guides instruction search using holistic evaluations or end-to-end task success. For complex multi-step reasoning, however, attributing a failure across a long trajectory remains difficult, as an argument may depart from a viable path long before its final line. Without fine-grained diagnostics, the optimizer must infer the cause of failure from sparse signals alone. We introduce Differential Prompt Optimization (DiffPO), a framework that supplies step-level diagnostic credit in natural language. During training, an auxiliary teacher model conducts a differential audit on failed student traces: it pinpoints the earliest step of divergence where the student's reasoning becomes unviable and abstracts that error into a generalizable rule for the prompt optimizer. To support the student at test time, the teacher also builds a library of general problem-solving strategies. The student selectively retrieves these strategies to guide its reasoning, providing high-level structure. Across different reasoning and instruction-following benchmarks, DiffPO consistently outperforms relevant prompt optimization baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.