ReTask: Compiling Execution Records into Tasks for Evaluation and Supervision
Abstract
Working agents routinely produce execution records, but reusable executable tasks for evaluation and supervision remain costly to construct and retain. An execution record provides only partial evidence about task inputs and completion conditions, mixed with results produced by one solver. We ask whether a single execution record can support a reusable task for evaluation and supervision without requiring exact recovery of the original task. \methodname first infers a task specification from the actionable request and evidence about the initial state. It then uses the specification to construct both an unsolved initial state for a new solver and an evaluator for the solver's outcome. Artifacts and outcomes produced during the recorded execution may identify content to withhold or provide validation cases for the evaluator, but they cannot define task requirements and are never shown to the new solver. For comparative evaluation, the observed Native and Compiled aggregate rankings agree on 20 of 21 pairwise orders among seven agent configurations on 100 paired Workspace-Bench tasks (). We further analyze two record conditions: the evaluated majority-synthesis and failed-source subsets contain solvable tasks with different model outcomes. Separately, for supervision, 2,040 new patches pass the original repository tests: 52.5% of all 3,889 Compiled attempts and 72.7% of the 2,807 evaluable patches. On the 1,123 tasks solved in both the Native and Compiled conditions, training on Compiled trajectories improves two base models on all six evaluated SWE-bench settings by 7.00–23.84 percentage points. These results show that tasks compiled from single execution records can support model comparison and new supervision without exactly reconstructing the original task.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.