acceptodds
Under review as a conference paper at ICLR 2027

Attribution Is Not Optimization: Counterfactual Rewards for Language Agents

Abstract

Holding a tool action fixed and replacing its output with a null observation isolates the output's contribution. It does not necessarily identify which action a policy should repeat. We formalize this attribution–optimization gap and introduce a finite-action audit that compares reward and terminal-utility update directions. Across three retrieval datasets, total belief gain is more aligned with terminal success than observation-residual credit (cosine 0.430 versus 0.114). The ordering persists in trainable-parameter gradients for retrieval (0.33 versus 0.09) and calculator use (0.51 versus 0.26). Yet residual credit closely tracks programmatic expression correctness (0.934). This contrast separates tool-output attribution from action learning. Restoring action-side credit improves local alignment throughout a controlled interpolation. We therefore introduce environment-marginalized action credit (EMAC), which preserves action-side credit while averaging tool-output variation. EMAC matches total belief gain when retrieval is deterministic. When the same query produces different relevant passage blocks, it raises held-out exact match from 0.263 to 0.312 and reduces normalized observation variance from 1.00 to 0.35. Under a matched reward-scoring budget, EMAC also improves exact match by 6.0 points in corrupted retrieval. On a 3B actor, the gain is 1.6 points.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.