Beyond Grouping: Orthogonal Advantage Decomposition for Critic-Free LLM Agent RL
Abstract
Group Relative Policy Optimization (GRPO) assigns the same rollout-level advantage to every step, which obscures credit assignment in multi-step LLM agents. Step-level methods such as GiGPO provide finer credit only when the same observation appears in multiple rollouts, but such exact matches are often sparse: in one ALFWorld reproduction, roughly one third of step groups had no match early in training. To move beyond exact grouping, we propose Orthogonal Advantage Decomposition Policy Optimization (OADPO), a critic-free method that replaces step grouping with a feature-based decomposition of expected return. OADPO shares evidence across nonidentical states through common feature components and assigns each component a closed-form shrinkage weight based on its estimated signal-to-noise ratio, allowing informative components to be trusted while noisy ones are suppressed. Because all steps in a rollout share the same outcome reward, OADPO corrects the noise estimate at the rollout level rather than treating steps as independent samples. Under independent features, the resulting weights are optimal within this estimator family and yield strictly lower prediction error than both GRPO's group mean and per-state averaging. The fitted state-dependent control variate is combined with a discounted step-level target, providing finer credit assignment without an additional critic, rollout, or model forward pass. Across ALFWorld, WebShop, and ScienceWorld with three model backbones, OADPO achieves the highest final success rate in seven of nine settings, including 96.4% on ALFWorld with Llama-3.2-3B, while matching GRPO's per-step compute in a WebShop timing check. An ALFWorld ablation further shows that the fitted control variate contributes a 6.0-point gain beyond the discounted step target alone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.