acceptodds
Under review as a conference paper at ICLR 2027

BILEVEL-BOOST: INTERLEAVING HYPERPARAMETER OPTIMIZATION WITH GRADIENT BOOSTING

Abstract

Gradient-boosted trees are powerful tabular predictors, but tuning them can require many expensive complete training runs. This cost can dominate model development and becomes especially burdensome in repeated experiments and rolling estimation. We replace repeated full-model search with HPO decisions made inside a single boosting trajectory. BILEVEL-BOOST grows alternative next blocks from a shared accepted prefix and uses validation feedback to adapt hyperparameters stage by stage, spending computation on local counterfactuals rather than discarded com- plete ensembles. This creates a tradeoff: smaller blocks increase control resolution but also increase the number of validation-dependent decisions. With tree budget T , block size b, HPO sample size nH , and at most q outcomes per intervention, we bound validation-selection error by O(pT log q/(bnH )). Combined with block- wise approximation error O((b/T )α), this yields a principled block-size scaling law. On a 100-seed drifting Friedman benchmark, BILEVEL-BOOST matches 8-trial Optuna in MSE using about 5× less wall-clock and 6× fewer trees; 40-trial Optuna improves MSE by only 1.06 at about 18× the wall-clock. Matched-trial open-loop schedules underperform, while block-size sweeps exhibit the predicted interior optimum and its qualitative shift with HPO sample size. In a rolling- window application to non-stationary financial data, BILEVEL-BOOST achieves roughly comparable predictive performance to a large-scale SOBER-tuned baseline using about 1/30 of the CPU-hours.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.