Gradient Boosting with Learnable and Incrementally Refined Trees
Abstract
Gradient boosting builds an ensemble by sequentially adding weak learners while keeping previously learned learners fixed. We revisit this design choice and ask whether boosting can benefit from jointly refining previously added learners as new learners are introduced. We investigate this question for gradient-boosted decision trees by replacing discontinuous split thresholds with differentiable routing functions, treating both tree decision boundaries and leaf values as continuously learnable parameters. Based on this formulation, we propose the Co-adaptive Gradient Boosting (CGB) framework and three algorithmic variants: CGB-AP, CGB-OB, and CGB-RO. All three methods perform gradient-based updates to the parameters of previously added trees. CGB-AP restricts boundary updates to axis-parallel directions, while CGB-OB permits oblique boundary updates. In both methods, newly added trees are initialized using standard purity-based greedy construction. CGB-RO instead bypasses greedy initialization entirely, starting each new tree randomly as an oblique decision tree and performing a localized gradient-based update before incorporating it into the ensemble. Across standard tabular benchmark datasets, we find that continuously updating previously added trees consistently matches or improves upon conventional discrete gradient boosting. These results suggest that freezing earlier weak learners is not a strict necessity of the boosting framework, and that incremental, differentiable refinement of the ensemble provides a promising alternative to the standard stagewise optimization paradigm.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.