acceptodds
Under review as a conference paper at ICLR 2027

DecisionOPD: Allocating Teacher Views and State Coverage for Decision Distillation

Abstract

Training compact language agents requires effective use of costly teacher supervision. Additional queries can refine targets for familiar decisions or extend supervision to new states, and the value of either choice depends on the training available to the student. We introduce DecisionOPD, a framework for comparing these allocations through executable action distributions while accounting for teacher cost and student exposure. The resulting comparisons connect supervision quality and breadth to downstream policy performance. Refined targets retain an empirical advantage over broader coverage as training increases, although the size of that advantage remains uncertain. On ALFWorld, aligned mean targets improve task success from 38.9% to 50.0% under matched state coverage and student updates.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.