acceptodds
Under review as a conference paper at ICLR 2027

Learning after Execution with Action-Specific Credit

Abstract

Execution is often used as terminal feedback: it scores a generated program, but the resulting state is not used for another learning action. Learning after execution instead uses the execution result as an intermediate state that supports an additional learning interaction for the same policy. This creates a credit-assignment problem because the two actions share the same final objective but play different roles in the trajectory. Accordingly, action-specific credit assigns the initial action credit from the final trajectory outcome and the corrective action credit for improvement over the intermediate state. Image-to-TikZ serves as an executable visual testbed for this setting. On 945 held-out examples, learning after execution strengthens the shared policy under single-call inference and also produces a post-execution corrective capability. Single-call render success reaches 92.70%, compared with 88.25% for continued single-stage GRPO after the same number of additional training iterations. This corrective capability is not reproduced by continued single-stage training and depends on the intermediate state. Alternative credit definitions also induce different group-relative preferences over the same trajectories. Together, these results suggest that executable tasks contain useful learning structure beyond a terminal reward.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.