Let Credit Follow Computation: Architecture-Aware Credit Transport for Large Language Model Reinforcement Learning
Abstract
Large-language-model reinforcement learning has substantially improved reward modeling, rollout sampling, baselines, and policy-update geometry, but it usually leaves one component architecture-agnostic: the operator that transports delayed evidence into token-level credit. Fixed-discount generalized advantage estimation (GAE) applies the same geometric kernel to every trajectory, while group-relative methods broadcast a response-level statistic to all tokens. We identify this discrepancy as an architecture–credit mismatch and introduce computation-conditioned credit transport (CCT), in which a detached statistic of the behavior policy's internal computation parameterizes a trajectory-specific causal kernel. Its concrete instantiation, CompPO, maps native attention concentration to a bounded retention gate, uses the gate in both the one-step bootstrap and a path-dependent advantage trace (Comp-GAE), and co-designs a M transport-aligned critic (TAC) that reuses actor hidden states and the policy KV cache. The external reward and clipped PPO objective remain unchanged, and a constant gate recovers fixed-coefficient GAE. Across five Qwen3-4B seeds, CompPO reaches final held-out accuracy (95% CI ), versus for GRPO from the same screen, under the same trajectory budget, with a best-to-final gap of versus points. Under that shared budget, CompPO reaches GRPO's peak at step . A controlled gate-by-critic factorial yields a -point final interaction ; shuffling and position-only controls show that trajectory-specific alignment contributes beyond mean scale and position. CompPO is stable in runs in a matched PPO stress grid versus for PPO, and improves frozen greedy macro accuracy over GRPO by and points on Qwen3-4B and Llama-3.1-8B-Instruct. These results establish policy-internal computation as a useful variable for RL credit estimation and motivate architecture-aware Bellman operators, traces, and critics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.