When Does Derivative Supervision Help? A Study of Label Construction and Noise
Abstract
Derivative labels describe local structure that value observations alone may not provide. Their usefulness, however, depends on how that structure is observed. We study label construction and loss balancing under a shared neural training procedure, then connect the results to controlled noise experiments. Across discontinuous-payoff problems, construction contributes more to value and derivative-error rankings than balancing, while the preferred construction reflects the information retained by the label. Varying value and derivative labels separately shows that the advantage of smoothed finite-difference (fuzzy) labels comes from their derivatives rather than from smoothing the value labels. Matched noise experiments reveal where useful derivative supervision becomes harmful to value prediction. A finite-design risk decomposition explains the trade-off between value-noise reduction, derivative-noise cost and approximation-bias change; its Fourier specialization shows how representation and sampling shape this balance. Combining a measured clean-label benefit with a last-layer estimate of the noise cost predicts whether derivative supervision still helps on new configurations for similar targets. Scientific extensions demonstrate value-prediction gains in a stochastic-volatility barrier problem and Burgers learning with a Fourier neural operator. Together with comparisons of higher-order labels, derivatives outside the value training region, and molecular objectives, these results establish a practical priority: determine what derivative information is supplied and how reliably it constrains the prediction before optimizing its weight.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.