Runtime Supervision for LLM Tutors Under Identical Prompting
Abstract
A correct answer is not always an appropriate tutoring response. Large language model (LLM) tutors must follow pedagogical requirements that depend on the learner’s prior attempts and misconceptions. Establishing the value of runtime supervision requires separating its contribution from prompting and checking whether LLM judges, models that read and score transcripts, recognize the failures being measured. We present a reference architecture and controlled evaluation of bounded, state-dependent supervision. The architecture checks conversation-dependent rules, such as withholding direct answers until a diagnostic question has been asked, and permits one corrective regeneration without substituting fallback responses; persistent violations are released and recorded. Across four nested conditions, the primary comparison uses the same tutor and identical initial prompts, with supervision enabled or disabled. We evaluate one proprietary and one open-weights model from a single developer on 20 data-science and statistics scenarios in 12 topic families, comprising 1,440 multi-turn dialogues and 800 single-turn stability probes. Assessment covers verifiability, stability, auditability, and pedagogical soundness under the Trusted Educational AI Standard (TEAS), alongside inference cost. Three LLM judges from different model families assess responses, with pedagogical judgments calibrated using openly licensed teaching exemplars and matched flawed alternatives. We test whether judges detect planted errors without flagging harmless changes; failure under prespecified criteria prevents confirmatory conclusions for the affected dimension. The pedagogy score represents the share of a scenario’s teaching requirements met, on a 0 to 3 scale. Instructions raise pedagogical adherence by 42% (+0.786 of 3, 95% interval [0.643, 0.929]). Adding reference context raises verifiability by 55% (+0.763 of 3) and moves auditability from 0.53 to 1.83 of 3. Supervision on top of the identical prompt adds 7% to pedagogy (+0.182 of 3, 95% interval [0.000, 0.364]), inconclusive under the registered rule. Delivery costs rise by factors of 1.39 and 4.14 for the proprietary and open-weights models, respectively, while released citation-check failures fall from 12.7% to 1.8% and from 32.4% to 3.4%. Stability falls by 17% with instructions and 20% for the full system; stability and auditability estimates are descriptive because their judges did not pass the planted-error examination.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.