Can LLM Examiners Explain Their Own Judgments?
Abstract
LLMs may replace or complement human decision makers in domains requiring professional discretion. In such settings, providing justifications for decisions may be ethically and legally necessary. LLMs can provide convincing justifications and explanations for their decisions, but it remains unclear whether these accounts correspond to decision policies expressed in their behavior. This issue is particularly important in high-stakes discretionary evaluation, where several considerations must be integrated without a fixed rule determining their relative importance. This study explores a pragmatic method to compare justifications and potential decision policies in the context of LLM-based exam grading using 120 synthetic answers to an authentic law exam. Each answer is scored independently by five LLM sub-judges on ten dimensions of answer quality derived from the exam grading rubric. The averaged sub-judge scores give an emulated objective quality assessment for each of the 120 answers. The examiner LLM, the main protagonist, assigns holistic grades from A to F and provides a justification for the same 120 answers without knowledge of the dimensions or the sub-judges' dimension scores. The examiner grades are then regressed on the sub-judge dimension assessments to examine which dimensions explain the variance in the examiner's grading most strongly. The resulting behavioral policy is then compared with a post-grading self-report and justifications given for individual grades. The study is replicated four times with different models as examiners (Gemma 3 (27B), Ministral 14B, DeepSeek V4 Pro, and GPT-5.6 Sol). The ten quality dimensions accounted for variation in individual grades (). The four examiner models largely agreed on which answers were stronger or weaker (), despite differences in grading severity. Behavioral cue-utilization profiles showed varying correspondence with both post-grading self-reports (–) and local justification profiles (–). Self-reports and local justifications were more consistently aligned with each other (–). The local justifications largely fit the predefined answer dimensions but also identified additional evaluative considerations. Overall, the findings suggest that LLM explanations can contain meaningful information about judgments, even without fully revealing how information is systematically weighted across decisions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.