SWPO: Score-Weighted Critic Fitting for Multi-Turn Policy Optimization
Abstract
As language agents generate longer reasoning and tool-use trajectories, learning from sparse outcome rewards becomes increasingly difficult. Critic-free group methods can receive weak or noisy learning signals in this setting, motivating PPO's learned state-dependent baseline. Turn-level PPO aligns optimization with complete reasoning-and-action decisions and can produce substantially shorter responses, but aggregating token scores and likelihood ratios can also increase update variability. Optimal-baseline theory motivates score-energy-weighted critic fitting to address this variability, yet existing OPT logit proxies require vocabulary-wide computation. We introduce SWPO, which constructs a cheap scalar proxy from stored sampled-token probabilities and sums it within each turn to weight critic fitting. The method retains the turn-level PPO actor and admits an exact factor-of-two logit-space bound, with a conditional interpretation of population baseline error. Experiments on search and embodied tasks support summed scalar weighting among the tested formulations, preserve shorter responses in the search-training comparison, and achieve competitive performance with the summed OPT proxy. SWPO reduces measured actor-update time by about 40% relative to summed OPT, with little additional time over the PPO baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.