KAIKA: Reinforcement Learning from Pairwise Comparisons of Continuations
Abstract
Reinforcement-learning post-training reduces a model’s behaviour to one scalar reward per response; Group Relative Policy Optimization (GRPO), the standard method, standardises these rewards within a group of responses to the same prompt. Many post-training outcomes are comparisons, a game won against an opponent or a puzzle solved where another attempt failed, and the scalar discards which continuation beat which. We ask whether a learner told only which of two continuations of the same state was better can post-train a model beyond GRPO. KAIKA is a population of policies that learns from such comparisons, with log-probabilities summed over each continuation: in games, a winner-vs-loser DPO term added to the policy gradient, pairing the played move with an alternative from the same position; on puzzles, where the comparisons replace the policy gradient, each successful attempt against each failed one. In chess, starting from a Stockfish-distilled 230M-parameter model, KAIKA gains +156 Elo over the base against GRPO’s +79 at T = 0.7 (paired score difference +0.0986, 95% CI [+0.0599, +0.1374], 5/5 seeds) and wins head-to-head on all 10 seeds; at T = 0.3 the paired difference is +0.0404 ([−0.0098, +0.0906]). The advantage over GRPO persists when every population member is an exact copy of the trained policy. On chess puzzles solved by 5M- and 32M-parameter language models, the comparison update exceeds GRPO with per-token-mean aggregation by +0.0098 at 5M (4/4 seeds) and +0.0127 at 32M (5/6 seeds), and by +0.0167 on three 32M seeds trained for 600 steps. Ablations locate part of the effect: replacing the summed log-probabilities with per-token means lowers accuracy on 3/3 seeds, by 0.0063 on average. Because an attempt ends at its first wrong move, failures are shorter than successes, and token averaging weights the correct opening moves a failure shares with a success more heavily in the failure; summing cancels their direct contribution. Neither population size (1 to 16 members), a curvature-based pair weight, nor a partial-credit reward adds measurable accuracy. Over training, the per-move rate of illegal moves rises by 41–91% under GRPO against 2–24% under KAIKA.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.