acceptodds
Under review as a conference paper at ICLR 2027

RecurSE: Recursive Self-Evaluation for LLM Rubric Judges

Abstract

LLM-as-judge is essential for evaluating open-ended text and steering post-training, yet improving the judge itself typically relies on expensive human annotations, auxiliary reward models, or distillation from stronger teachers. In this work, we consider that judging another model's response and judging one's own evaluative reasoning are essentially the same capability. We therefore improve the judge model's evaluative capability without expensive external gold supervision: the model's own evaluations generate the learning signal in a closed-loop regime of recursive self-improvement (RSI) that we term Recursive Self-Evaluation (). Specifically, employs two-pass reinforcement learning: a trainable judge model evaluates candidate responses under per-rule rubrics (Pass 1), while a synchronized copy of the current policy (the checker) audits the judge's reasoning against meta-rubrics to supply a scalar process reward (Pass 2). We then investigate two governing questions of this loop: when can self-improvement occur, and when must it stop? For the first question, we make learning viable via interface decoupling, which structurally isolates the checker's scalar score from the judge's verdict tokens and closes a degenerative token-copying shortcut that inflates self-assigned rewards. For the second, because unanchored recursive learning is inherently bounded, we introduce Pairwise Advantage Validity (PAV), a validation monitor that jointly tracks judge accuracy and checker fidelity to reliably identify a useful early-stopping window. Across Qwen3.5-9B, Gemma-4-E4B-it, and Qwen3.6-27B, achieves consistent generalization gains across held-out medical, pairwise, summarization, and professional benchmarks. Extensive ablations show that synchronized judge–checker co-evolution outperforms frozen checkers, external meta-judges, self-consistency, and scaled teacher distillation. Preference pairs curated by our judge model further improve downstream policy alignment. RSI for LLM-as-judge is thus feasible when self-produced reward validity is explicitly decoupled and monitored.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.