Categorical Regression in REBEL: Rate Preservation, Boundary Bias, and No General Robustness
Abstract
REBEL reduces policy improvement to squared-loss regression of reward gaps onto policy log-ratio differences and asks whether cross-entropy does better. We answer for the tied HL-Gauss loss, which encodes the scalar as a cross-entropy-trained histogram. Inside the support the loss is two-sidedly calibrated to the squared loss with constants free of support width and bin count, so REBEL's bound carries over, with its square-root rate sharp under the stated assumption class. At the boundary the encoding winsorizes noisy targets, a bias absent from squared regression, bounded by how far and how often targets cross the boundary. Preregistered bandit experiments confirm this picture: under symmetric contamination HL trails MSE in all 20 seeds, by 3.4% of clean regret on average. One-sided, feature-correlated spikes give a replicated benefit of 0.7% of clean regret near the interpolation threshold, below the preregistered smallest effect of interest and, in the registered runs, not separated from curvature-matched Huber. Widening the support attenuates both effects; at LLM scale HL and Huber are equivalent within the registered margin of ±0.035 clean-reward units. The replacement keeps REBEL's rate but buys no general robustness: what matters is where a loss clamps, not whether it is categorical; curvature-matched Huber remains the conservative default.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.