acceptodds
Under review as a conference paper at ICLR 2027

Fifty Rewards, One Update: Auditing Objective Capacity Beyond Mechanism Activation

Abstract

Multi-reward language-model optimizers are commonly described by the number of rewards they expose, but specified rules, aggregation tiers, emitted channels, empirical advantage rank, objective-gradient span, and capacity consumed by an optimizer are different quantities. We introduce an objective-capacity audit that traces these quantities from reward definitions to parameter updates. Pointwise within a shared rollout batch, every fixed linear combination of per-objective policy gradients is exactly equivalent to one scalar advantage. ObjCapBench shows that aligned controls collapse, structural panels activate PCGrad and resist held-out reconstruction by one batch-invariant weighting, and natural activation is rare and model-dependent. A fixed-rank intervention passes 31 of 35 mechanism and prediction gates, while a synthetic construction preserves nominal width 50, rank, mean, contrast energy, and the full singular spectrum yet changes PCGrad under an objective-basis rotation. Crucially, a separately sealed two-model, five-seed downstream test finds no positive advantage over the paired per-cell maximum of three registered static challengers: the estimate is , with a 95% interval , and the preregistered primary gate fails closed. Nine of ten blocks lie at the endpoint's zero floor, limiting sensitivity but not changing the registered decision. Mechanism activation and non-reducibility therefore do not imply downstream benefit; no universal scaling law or algorithmic-superiority claim is supported.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.