acceptodds
Under review as a conference paper at ICLR 2027

Learning Before the First Success: Bidirectional Progress-Guided Training from Failed Cyber Interactions

Abstract

Training large language model (LLM) agents on difficult interactive tasks faces a bootstrapping barrier: when no rollout succeeds, terminal rewards make distinct failures indistinguishable. We study how failed interactions can still drive learning before the first complete solution. We introduce Bi-PGSD, a bidirectional progress-guided framework that pairs execution-verified forward effects with backward goal reasoning to construct locally checkable objectives. These objectives provide environment-anchored rewards for policy updates; target-side validation enables cross-task reuse; and, after a complete path is independently reproduced, retrieval and intermediate guidance are progressively withdrawn to consolidate autonomous behavior. Under a fixed computation budget, Qwen3.5-9B on Cybench strict zero-success (Z) improves first-success discovery from 50.0% with strong-summary retrieval to 62.5%. On scaffold-free Cybench family-held-out tasks, the trained policy improves complete-task success by 10.3 percentage points over the initial policy. Across mechanism and transfer controls, the results show that failed trajectories become useful training signals when local effects are both verified and directed toward the terminal goal.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.