Online Alignment of Ranking Model via Variance-Weighted Multi-Generator Feedback
Abstract
We study online alignment of a pointwise ranker to the feedback of multiple downstream generators in retrieval-augmented generation (RAG). Existing work usually aligns the ranker to a single generator, which does not reflect the common deployment reality that one ranker serves several downstream models of different capabilities. When extending to scenarios involving multiple generators, how to aggregate the rewards from multiple generators becomes a key issue. A naive fix, averaging the utilities of all generators uniformly, dilutes the learning signal, because it gives equal weight to generators that are largely insensitive to document order. We propose Variance-Weighted Multi-Generator GRPO (VW-MG-GRPO). For each query we draw G top-K permutations with a Plackett–Luce policy, send them in parallel to a heterogeneous pool of generators (weak / mid / strong), and aggregate their utilities using each generator's within-group standard deviation as its weight. On Natural Questions, TriviaQA, and HotpotQA we validate the method and, more importantly, give a systematic analysis of the variance dynamics during training. We report three transferable findings: (i) variance weighting acts as an implicit per-query router rather than a smoother; (ii) the total variance signal decays over training, so variance weighting behaves as a self-annealing curriculum; and (iii) the effective per-query signal is that of uniform aggregation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.