BEYOND GROUP SIZE : THE SPECTRAL GEOMETRY OF PLACKETT–LUCE SUPERVISION
Abstract
LLM post-training increasingly uses ranked groups of candidate responses rather than isolated pairwise choices, raising the question of how much local Fisher information one complete-ranking observation carries about the underlying relative scores. A complete ranking is not equivalent to observing every pair independently: the Plackett–Luce model generates it sequentially, so a pair contributes only while both members remain unselected. We show that the exact Fisher operator of a complete ranking is a weighted graph on the candidates whose edge weights factor into two logically distinct terms. The first is classical Bradley–Terry discriminability, fixed by the pair's own score gap relative to the evaluator's resolution. The second we call exposure: it records how the sequential process attenuates that information, and it changes when the other candidates move even though the pair's own scores do not. We prove exposure never exceeds one and never falls below one part in the group size minus one. The exact operator is therefore sandwiched, in the positive semidefinite sense, between two multiples of the Laplacian of a reward-gap graph whose edges are heavy between candidates that are hard to tell apart, so the weakest identifiable Fisher curvature of a full ranking, the relative-score contrast it constrains least, is pinned by that graph's algebraic connectivity to within a factor of the group size minus one. Symmetry calibrates this picture rather than explaining it: there the curvature has a closed form, rising with group size but saturating while ranking entropy and total Fisher mass keep growing, and stagewise curvatures provably cannot be added. Group size and reward range therefore do not determine this curvature, and we test total Fisher mass empirically. Across 9,696 exact instances the bounds hold without genuine violation; at fixed group size and range the curvature moves by a factor of 8384 while the trace moves the other way, and nearly trace-matched configurations differ hundreds-fold. These are local information-geometric results, not finite-sample claims.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.