Solved ≠ Learned: When Outcome Optimization Rewards Answer Disclosure in LLM Tutors
Abstract
We study outcome optimization in open-ended LLM tutoring, where learner improvement is not directly verifiable and must be estimated from dialogue or simulation. In this setting, the tutor can also influence the evidence used to measure progress by revealing solution-relevant information. We refer to this as disclosure inflation. We adapt three representative state-of-the-art LLM tutor training objectives—pedagogical-process, turn-level outcome, and trajectory-level outcome optimization on open-ended coursework-based tutor–student dialogues while keeping the training setup fixed. The Pedagogically Optimized Tutor achieves the highest long-horizon KC mastery (0.85), compared with trajectory-level (0.81) and turn-level outcome optimization (0.72), despite using no simulated look-ahead or trajectory-level mastery reward. It also shows the lowest answer disclosure in goal-oriented evaluation. Trajectory-level planning improves over turn-level optimization, but does not outperform the process objective. On a separate verifiable mathematics evaluation, tutoring improves accuracy from 14.7% to 26.7%. These results show that optimizing a model-estimated learner outcome does not necessarily optimize independently measured learner mastery.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.