Can LLMs Judge Better Than They Generate? Task Asymmetry, Mechanistic Interpretability and Transferability for In-Context QA
Abstract
LLM-as-a-Judge, self-correction, and RLHF reward models all assume that judging an answer is easier than producing one. This assumption is rarely tested directly. We test it in a controlled in-context QA setting where a single passage is the only source of information and each model judges the answer it just generated, so any difference between generating and judging comes from task difficulty, not memorized facts. Closed-book and counterfactual controls support this, because removing the passage lowers generation accuracy by 44 to 65 points. We find that judging is not consistently easier. Across four models (Llama-3.1-8B, Qwen3-8B, Qwen3-32B, GPT-4o-mini) and four benchmarks (SQuAD 2.0, DROP, HotpotQA, MuSiQue), the sign of the gap changes with the model and the dataset. One mechanism underlies this variation. The evaluator does not re-read the passage. Last-token attention shows it attends very little to the context or to the candidate answer, and a causal ablation shows its verdict depends on the candidate it is given, not on a fresh reading of the passage. Evaluation is therefore a shallow check. The same mechanism explains both directions of the gap and is hard to remove: it remains when the judge reasons step by step, it grows with model size as larger judges catch fewer of their own errors, and it survives LoRA fine-tuning, which shifts the asymmetry without removing it. Fine-tuning on generation alone makes a model accept almost every answer when it is later reused as a judge (recall near 100%). These self-evaluation failures come partly from how attention is used at inference time, a practical concern for pipelines that let a model judge its own outputs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.