When Does Curvature Help Offline Value Selection?
Abstract
Repeated Bellman updates fit the same finite dataset more closely, but need not improve the true value estimate. We study risk calibration and estimator selection with fixed linear features. Joint reward, transition, and Gram-matrix uncertainty changes both the variance and the mean of a finite iterate; omitting the mean correction can miss an order- risk term. We derive a same-data risk score for smooth value estimators and extend its guarantees to ordinary independent and identically distributed (iid) transitions through local conditioning and an explicit fallback on ill-conditioned datasets. For one-state problems, a characterization for any sample size and finite candidate bank identifies when the correction's downward shift improves selection and when it hurts. We also show why an endpoint-scale oracle bound alone cannot certify improvement over convergence. Matched experiments evaluate 19,800 independent datasets, including 7,200 iid datasets with overlapping features, using matched ridge, normalized-difference tournaments, paired intervals, and independent pilot comparators for each candidate family. Curvature gives modest, statistically resolved gains within both checkpoints and ridge, but independent pilots reveal much larger remaining selection gaps. Severe small-sample miscalibration and the cost of covariance estimation remain important limits. We report guards that reject the population and conservative confidence checks that certify no fits; neither calibration nor useful selection follows from an observed inverse margin alone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.