CRAFT-GBM: Selecting Best Features for Gradient-Boosted Machine Learning
Abstract
Feature subsampling is widely used to accelerate gradient-boosted trees, yet standard approaches choose the number of features in advance and largely ignore the split-gain evidence accumulated during training. We introduce CRAFT-GBM (Confidence-Ranked Adaptive Feature Tracking), a native LightGBM extension that uses this evidence to adapt which—and, by default, how many—features are searched in each boosting round. CRAFT-GBM transforms root-split gains into bounded observations and constructs joint time-uniform bounds that remain valid under adaptive, non-stationary training histories. In its default endogenous mode, after initialization, CRAFT-GBM retains exactly the features that cannot yet be ruled out as best, allowing the selected feature-set size to emerge from the observed training history rather than from a pre-specified fraction. With probability at least , every feature with the largest expected gain is retained jointly across all rounds after initialization. When a feature fraction is specified instead, CRAFT-GBM ranks features by their interval point estimates and provides an approximate top- guarantee for the corresponding feature budget . We implement CRAFT-GBM directly within LightGBM while preserving its native feature-bundling representation, and evaluate it across 14 public and synthetic dataset families together with four noise-augmented settings. In the fully endogenous setting, CRAFT-GBM reduces per-fit training time by 39.74% while achieving 1.04% lower test loss in the geometric mean across all 18 RQ1 experiments. Under pre-specified feature fraction , Exogenous CRAFT-GBM also yields aggregate training-time reductions up to 13.52 percentage points larger than random feature subsampling, with no increase in aggregate test loss relative to full-feature LightGBM. These results show that accumulated training evidence can support adaptive feature search with statistical guarantees while retaining the practical advantages of native gradient boosting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.