acceptodds
Under review as a conference paper at ICLR 2027

Signal Reshaping for GRPO in Weak-Feedback Agentic Code Repair

Abstract

Code-agent RL often operates under weak online execution feedback: rollout-time signals are reliable and executable, but capture only necessary conditions rather than the target semantic predicate. Using agentic compile-fix as the setting, we study signal reshaping for GRPO with offline reference-conditioned reward construction. Our central claim is that GRPO's within-group comparison is meaningful only after three kinds of signals are reshaped: outcome rewards recover semantic ranking, process signals localize intra-trajectory credit, and rollouts from the same prompt remain execution-comparable. We operationalize these conditions by reshaping signals around GRPO: compile-and-semantic layered rewards reshape trajectory ranking, step-level process scores placed outside group reward normalization act as token loss weights to reshape intra-trajectory update strength, and failure-cause-aware rollout governance preserves within-group comparability. Experiments show a clear end-to-end gain: full signal-reshaped GRPO improves strict compile-and-semantic accuracy from the base model's zero-shot to . Controlled comparisons further explain the source of this gain: binary rewards remove the compile-only middle tier and degrade trajectory control; on top of layered rewards, process-score weighting further improves accuracy from to and reduces average evaluation steps from to . As a boundary comparison, token-level privileged self-distillation underperforms step-level process weighting, showing that local distribution matching cannot replace action-step credit.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.