When Should an AI Tutor Withhold Feedback? Selective Instruction under Diagnostic Uncertainty
Abstract
Effective pedagogical AI systems must guide students to correct errors without prematurely revealing solutions. However, this objective is challenging when an error stems from multiple plausible causes, where even a pedagogically accurate response risks inadvertently leaking the answer. To address this, we propose PARS-Tutor, a tutoring policy that maintains multiple candidate error hypotheses and systematically suppresses feedback that fails pedagogical or safety constraints. Evaluated on 100 held-out mathematics problems, PARS-Tutor with a budget of seven hypotheses achieved a 64% success rate in generating valid, non-revealing feedback across all scheduled cases (versus 60% for direct prompting and 59% for chain-of-thought prompting); among delivered responses, its success rate reached 67.4% (versus 62.1% and 60.0%, respectively), with zero detected answer leakage across all 71 returned outputs. Furthermore, scaling the hypothesis budget from five to seven improved candidate availability by 27 percentage points, explicitly highlighting the trade-off between safety and abstention. Our findings demonstrate that post-hoc verification formalizes the tension between safety and response availability, providing a grounded framework for reliable AI tutoring while pointing to future work on downstream learning gains.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.