Dual-Axis Policy Optimization for LLM Agents: Bayesian Feedback Attribution and Trajectory Mass Normalization
Abstract
A reinforcement-learning objective for LLM agents makes two weighting choices that are usually left implicit: which decisions inside a trajectory receive credit, and how complete trajectories are weighted in the batch. We make both choices explicit. For the first, we propose Bayesian Feedback Attribution (BFA): we score the observed environment feedback under the executed action and under a counterfactual action sampled from the same policy, and use the resulting posterior lift to reallocate the host learner’s advantages across decisions while conserving each trajectory’s total coefficient mass. The policy model itself serves as the feedback model, so no learned dynamics model, return predictor, or auxiliary loss is required. For the second, we show that global token averaging weights trajectories by a size-biased transform of the episode distribution, and we derive the exact finite-batch discrepancy to uniform per-trajectory averaging as a length–gradient covariance term. Trajectory Mass Normalization (TMN) weights episodes uniformly, matching how benchmarks evaluate: one vote per episode. The resulting framework, BATON, introduces no new hyperparameters and no additional environment interaction, and drops into GRPO and GiGPO. On ALFWorld, WebShop, and search-augmented question answering across 1.5B–7B models, each component improves its host on its own, and the combination is best or tied for best in every comparison we ran, at 10–17% additional wall-clock cost measured on ALFWorld at the 1.5B scale.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.