acceptodds
Under review as a conference paper at ICLR 2027

Decomposed Prompt Optimization For Reliable LLM-as-Judge

Abstract

Generative LLM-as-judge systems for safety guardrails, reward models, automatic evaluators, and niche domains are typically configured by a natural-language prompt - a flexible lever for making a weak judge stronger without modifying its weights. We study how far this lever goes for small frozen judges, and find that unconstrained prompt optimization i.e. rewriting the prompt as a single free-text block, buys surprisingly little. The missing ingredient in here is *decomposition*: a judge prompt splits into two functionally distinct components: *criterion* (what condition to detect), and *scoring schema* (how to map a judgment to a decision), and making this decomposition *explicit* to the optimizer is what makes automatic prompt optimization pay off. Across four judge models (B-B) and six tasks we find that (i) reflective optimization improves every judge, most where it is weakest (up to F1); (ii) treating the prompt as two explicit components beats optimizing it as one monolithic block on every judge (up to F1); and (iii) *which* component to search and in what order, is the most effective way to *spend* a compute budget: phasing converts extra search into the largest gains on the hardest structured tasks (up to F1) once initial search saturates. Applying this recipe to a real-world financial-compliance policy that fixed-taxonomy guards cannot express, a frozen B judge, optimized only through its prompt, improves F1 from to while baseline guards stall at -.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.