acceptodds
Under review as a conference paper at ICLR 2027

Solved ≠ Learned: When Outcome Optimization Rewards Answer Disclosure in LLM Tutors

Abstract

We study outcome optimization in open-ended LLM tutoring, where learner improvement is not directly verifiable and must be estimated from dialogue or simulation. In this setting, the tutor can also influence the evidence used to measure progress by revealing solution-relevant information. We refer to this as disclosure inflation. We adapt three representative state-of-the-art LLM tutor training objectives—pedagogical-process, turn-level outcome, and trajectory-level outcome optimization on open-ended coursework-based tutor–student dialogues while keeping the training setup fixed. The Pedagogically Optimized Tutor achieves the highest long-horizon KC mastery (0.85), compared with trajectory-level (0.81) and turn-level outcome optimization (0.72), despite using no simulated look-ahead or trajectory-level mastery reward. It also shows the lowest answer disclosure in goal-oriented evaluation. Trajectory-level planning improves over turn-level optimization, but does not outperform the process objective. On a separate verifiable mathematics evaluation, tutoring improves accuracy from 14.7% to 26.7%. These results show that optimizing a model-estimated learner outcome does not necessarily optimize independently measured learner mastery.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.