acceptodds
Under review as a conference paper at ICLR 2027

StepWeb: Trajectory-Conserving Step-Level Credit for Multi-Turn Web Refinement

Abstract

Multi-turn webpage refinement requires agents to improve their own implementations while operating on states shaped by earlier edits. Terminal rewards can identify successful trajectories, but provide limited information about which intermediate decisions caused regression, recovered lost quality, or established genuine progress. We introduce STEPWEB, a trajectory-aware reinforcement learning framework that separates terminal trajectory comparison from within-trajectory credit assignment. Its trajectory-conserving step-level credit (TC-SLC) assigns history-aware, magnitude-sensitive signals to Regression, Repair, and New Gain events, and redistributes them across refinement turns through token-length-weighted centering while preserving the trajectory’s mean advantage before policy clipping. Experiments on Vision2Web Level 1, Design2Code, Design2Code-HARD, and Interaction2Code show consistent improvements over the Qwen3.5-9B backbone, yielding a 16.4% macro-averaged relative gain across the four benchmarks. In controlled multi-turn refinement on Vision2Web Level 1, TC-SLC further reduces the edit regression rate by 29.5% relative to terminal-only RL while improving final reconstruction quality. These results demonstrate the effectiveness of trajectory-local credit assignment for reliable iterative visual coding.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.