acceptodds
Under review as a conference paper at ICLR 2027

Voting with the Graph: Consensus Rewards for RLAIF via Topological Consistency

Abstract

Pairwise LLM judgments provide local preferences, whereas group-based policy optimization requires scalar rewards. We study how to aggregate these comparisons for Reinforcement Learning from AI Feedback (RLAIF). We introduce Topological Consensus Rewards (TCR), which uses a greedy feedback-arc-set heuristic to retain an acyclic subset of comparisons and computes net-degree rewards on the retained graph. Its distinguishing operation is evidence selection: raw win-tie rates and original-graph net degrees are equivalent under group standardization, up to the numerical stabilizer. We use Cycle Incidence Rate (CIR) as a structural diagnostic and analyze triangle participation under an independent-error model to motivate the aggregation choice. Experiments on Arena-Hard, MT-Bench, and WritingBench evaluate TCR in two Qwen3 training settings, with higher mean Arena-Hard overall scores than the compared pairwise baselines. The results support transitivity as a useful reward-aggregation bias when a single preference ordering is an appropriate approximation, rather than a guarantee of preference-error recovery.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.