Beyond Final Regret: Stress-Testing Hyperparameter Optimization Conclusions Across Evaluation Regimes
Abstract
Hyperparameter optimization (HPO) methods are commonly compared on fixed collections of tasks using aggregate final performance, implicitly treating the benchmark and evaluation protocol as fixed. We study a different question: how sensitive are empirical conclusions about HPO methods to controlled changes in the evaluation regime? We introduce Beyond Final Regret, a configurable benchmark generator that exposes search-space structure, fidelity reliability, cost regime, distribution shift, evaluation protocol, and backend realization as explicit experimental variables. The framework generates reproducible benchmark instances and retains objective, fidelity, incumbent, and cost trajectories, enabling evaluation through final quality, anytime behavior, quality–cost trade-offs, conditional comparisons, and repeated-run distributions. We evaluate seven representative HPO methods across in-distribution, interpolation, extrapolation, cross-family, and targeted stress-test settings. The experiments show that several conclusions obtained from aggregate final performance are conditional on the evaluation regime. Relative performance changes with search-space structure, fidelity reliability, available resources, distribution shift, protocol, and backend, while other relationships remain stable across perturbations. In particular, methods with strong final quality can incur substantially larger computational costs, making recommendations dependent on the available resource budget rather than on a single cost-normalized scalar. These results position benchmark design itself as an experimental variable and provide an executable framework for measuring the sensitivity of HPO conclusions to that variable.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.