When Counterexamples Don't Teach: Learner-Specific Feedback in LLM Repair
Abstract
Counterexamples are widely used to guide program synthesis and repair: a verifier exposes a concrete failure, and the learner revises its current solution from that feedback. Existing selection strategies largely judge counterexamples by how informative they are about the hypothesis space. This is natural for symbolic learners, but less obvious when the updater is an LLM, whose response to the same valid feedback can vary substantially. We show that formal informativeness is often a poor predictor of whether an LLM will make a useful repair. The mismatch appears across several LLMs and executable task families. Motivated by this observation, we introduce TEACHCE, which selects counterexamples according to the repair consequences they are expected to induce in the current learner, and an adaptive verifier-guided variant that reuses repair attempts when needed. Across controlled repair tasks and sequential settings, this learner-specific view improves the quality–compute trade-off over symbolic selection. On public student-program benchmarks, we further show that combining consequence-aware selection with adaptive verifier-guided rollout reuse improves repair efficiency under a limited inference budget. These results suggest that, once LLMs enter the repair loop, choosing feedback requires modeling not only what the evidence says, but also how the learner will respond to it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.