acceptodds
Under review as a conference paper at ICLR 2027

Same Groups, Different Values: Decomposing Step-Level Credit Assignment for LLM Agents

Abstract

Step-level group-relative estimators for language-model agents are often characterized by the states they place in the same group. We show that this description is incomplete: the rollout partition determines where statistics are shared, whereas the value functional determines what signal is aggregated. We introduce a component-resolved audit that varies these choices separately on matched rollouts, and complement it with a support-aware variance decomposition for measuring history-associated value variation after the task and current observation are fixed. The framework combines an unbalanced random-effects estimator with task-level resampling, explicit accounting of variance-bearing groups, and replay-based validation of the collected trajectories. Across replay-verified policy snapshots, two text environments, and an independently emitted matched-row batch, the audit reveals that apparently different grouping methods can remain closely aligned when they retain the same base operator, while added value machinery can dominate their differences. It also shows why sparse terminal feedback can leave most formally eligible groups uninformative even when a denser progress signal varies within them. The academic value is a reusable attribution protocol: it separates algorithmic structure, measurable support, inferential units, and data fidelity before a snapshot difference is assigned a mechanism.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.