Teaching Is More Than Solving: Compiling Executable Lessons for Agent-to-Agent Transfer
Abstract
Language-agent systems routinely ask one agent to summarize experience for another, yet they evaluate the sender as a solver rather than as a teacher. We separate these capabilities by defining teaching utility: the improvement that a budgeted lesson causes in a frozen student's decisions on disjoint held-out states. We introduce TeachWorld, a controlled testbed of 144 latent decision policies spanning six semantic domains, three rule structures, and 124 distinct family/structure signatures, and evaluate 432 teacher-concept artifacts across three independently served model families. We then present LessonSmith, a training-free lesson compiler that pairs a teacher's candidate rule with branch-covering executions from an exact verifier and makes every confirmation or correction explicit. Agent-written lessons improve frozen students by 44.6 percentage points over no lesson. Given an executable verifier and a branch-coverage mechanism, LessonSmith adds 4.3 points (95% concept-clustered CI [2.8, 5.9]). Two full-matrix controls identify why. Given the same executed evidence, explicit receipts outperform an execution-informed rule rewrite by 2.7 points ([1.2, 4.2]). When teachers instead choose four probes from eight fresh unlabeled cases, they cover only 2.40 of four policy branches and trail branch-covered lessons by 6.6 points ([4.8, 8.3]). A 12-item solving battery predicts teaching () but still misorders 27.5% of comparable teacher pairs. Teaching is therefore a measurable transfer capability: solving quality matters, while executable coverage and the form in which evidence is communicated determine what the next agent can use.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.