Why and When Score-Function Learning Fails under Discrete Downstream Evaluation
Abstract
Score-function learning estimates gradients from scalar feedback and thereby enables optimization through nondifferentiable and even black-box evaluators. Although its estimator is unbiased for a smoothed objective, this guarantee does not determine whether the current gradient can be recovered under a finite evaluation budget. We identify two failure mechanisms under discrete downstream evaluation. Boundary starvation occurs when sampled candidates rarely cross loss-changing decision interfaces, causing centered group updates to vanish. Tangential noise arises because perturbation components tangent to these interfaces contribute variance without changing the mean gradient. Based on this characterization, we introduce two budget-dependent quantities. Activation evidence measures whether the available queries are likely to reveal an informative loss comparison, while gradient resolvability compares the squared norm of the smoothed gradient with the estimation variance remaining after budgeted averaging. Both quantities can be estimated from pilot groups through a confidence-aware diagnosis. Numerical experiments verify the existence of both failure mechanisms and the effectiveness of the proposed quantities.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.