Under Observation: Mitigating Review-Induced Bias in LLM Reflection
Abstract
Self-reflection serves as a critical capability for large language models to detect errors, revise prior outputs, and ensure reliable agentic operations. To enhance this process, many evaluation frameworks and practical agent systems explicitly inform the model that its output is being reviewed, assuming this evaluative context will encourage more rigorous reassessment. However, we demonstrate that these review cues fail to improve overall performance and instead induce a systematic bias where models persistently maintain their initial answers. Further analysis reveals that this bias stems from a stable review-induced state shift that causally alters model judgments, even when accurate correctness information remains preserved within internal representations. To address this issue, we propose ReSET, a Review-Shift Editing Training framework designed to restore reliable self-reassessment. ReSET combines correctness-verified teacher responses with answer-distribution distillation to train models to re-solve problems before making judgments and effectively neutralize the review-induced shift. Comprehensive evaluations demonstrate that ReSET consistently reduces wrong-answer persistence while preserving correct-answer retention, leading to improved judgment discrimination and strong cross-task generalization. Ultimately, this work reveals the inherent behavioral shift of large language models under external observation, offering valuable insights for the future optimization of reliable LLM self-evaluation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.