Learning Reusable Reasoning Critics from Expert-Referenced Supervision
Abstract
Process-level reward models can score intermediate reasoning, but they typically require step-level annotations, preference data, or external verifiers. We study whether paired expert demonstrations can instead train a reusable reasoning critic without any of these. A naive expert-versus-policy classifier risks learning source or formatting cues instead of reasoning quality. We introduce the **Prefix-level Expert-Referenced Critic (PERC)**, which avoids this by using each expert trace three ways: as a positive example, as the reference answer for weakly labelling policy rollouts by agreement, and as the basis for targeted corruptions used as hard negatives, so that policy traces appear on both sides of the classifier and provenance alone cannot separate positives from negatives. Across GSM8K, MMLU-Pro, and MedReason, this critic captures a reusable reasoning-quality signal. As an **inference-time ranker**, it improves Best-of-16 pass@1 in all main reranking settings, transfers positively across all 27 source–target pairs with gains up to 12.7 points, complements consistency-based voting, and continues to improve pass@1 even for policies already trained with it. As a **process-level evaluator**, prefix-value changes localise human-labelled MedReason error units better than chance and entropy baselines and identify controlled GSM8K perturbations with up to Hit@1. It is also a viable **training signal** for GRPO and remains competitive with imitation baselines without task-specific reward tuning. Overall, PERC shows that weak expert-referenced supervision yields reusable critics for inference-time selection and error diagnosis, while remaining feasible as a policy-training reward.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.