acceptodds
Under review as a conference paper at ICLR 2027

biCurriculum: Crediting Executable Task and Environment Rewrites by Their Learning Utility

Abstract

A failed agent trajectory identifies where a policy breaks, but not which training intervention will improve the unchanged task. We introduce biCurriculum, a framework that turns failures into executable task or environment rewrites and credits them by downstream learning utility relative to matched original-task replay. The framework separates native rewrite validity, solver reachability on the child, and original-task improvement after training. On complete WebArena-Lite165, a Qwen3.5-9B rewritten-data arm solves 44/165 tasks versus 34/165 for a replay control and 29/165 for the starting solver; the paired run has 13 rewrite-only versus 3 replay-only successes (exact McNemar ). Across three completed training runs, the rewritten-data arm solves 43–44 tasks while replay solves 33–35, a mean gain of 5.9 percentage points. Controlled ALFWorld results are heterogeneous across the pre-specified matched batches: two matched batches score 107/128 versus 118/128 and 108/128 versus 110/128, while a third reaches 120/128 versus 109/128 (13 versus 2 discordant successes, ). A diagnostic audit shows that only 2/10 originally trained children exercise the parent failure step. On the same three B1 training parents, a failure-step-preserving correction raises native sampled success from 2/12 for the original curriculum to 12/12 (start 5/12; replay 10/12). Together, the results show both when executable rewrites can beat replay and why valid, reachable rewrites can fail: the child must expose the capability missing on the parent.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.