acceptodds
Under review as a conference paper at ICLR 2027

Deviation Is Not Failure: Decoupling Plan-Adherence from Task-Correctness in LLM Judges via Input-Isolated Single-Axis Fine-Tuning

Abstract

As agents execute longer tool-rich workflows, evaluating them increasingly relies on an LLM judge that inspects the execution trace. Such a judge is implicitly asked two orthogonal questions: did the agent deviate from its stated plan (a procedural question, DEV), and did the outcome violate the task objective (a correctness question, VIOL)? We first show that state-of-the-art judges systematically conflate these two axes: on tau-bench traces whose two gold labels disagree, even the strongest API judge with a carefully decoupled, evidence-forced prompt returns the same verdict for both axes 87–90% of the time and answers both correctly only 9–12% of the time–and an explicit "decouple-then-judge" prompt is worse than a joint one. We trace the cause to a shared input: a single judge that sees the plan, the spec, and the trace at once cannot help but let procedure color its correctness verdict. DecoupledJudge turns this diagnosis into a method, training two single-axis judges whose inputs are isolated by construction–the DEV judge sees only the plan and the trace, the VIOL judge sees only the spec, the trace, and the delivered outputs–so neither axis can copy the other. On the orthogonal-conflict subset, DecoupledJudge cuts conflation from 0.71 (untuned base) and 0.54 (best API prompt) to 0.16, and raises the fraction of traces on which both axes are correct from 0.17 to 0.84, with paired bootstrap CIs excluding zero and an exact McNemar test at p<10^-12. The decoupling transfers under leave-one-domain-out, a controlled ablation shows the input isolation itself (not the training data) is causally responsible, and the effect is scale-invariant across 3B, 7B, and 14B, confirming that conflation is an architectural artifact of shared inputs rather than a small-model deficiency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.