DecisionOPD: Allocating Teacher Views and State Coverage for Decision Distillation
Abstract
Training compact language agents requires effective use of costly teacher supervision. Additional queries can refine targets for familiar decisions or extend supervision to new states, and the value of either choice depends on the training available to the student. We introduce DecisionOPD, a framework for comparing these allocations through executable action distributions while accounting for teacher cost and student exposure. The resulting comparisons connect supervision quality and breadth to downstream policy performance. Refined targets retain an empirical advantage over broader coverage as training increases, although the size of that advantage remains uncertain. On ALFWorld, aligned mean targets improve task success from 38.9% to 50.0% under matched state coverage and student updates.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.