acceptodds
Under review as a conference paper at ICLR 2027

CAST: Learning Compact Decision Trees via Compound Additive Splits

Abstract

Decision trees expose their routing rules, but simple splits can require large trees to represent nonlinear boundaries. Nonlinear splits can form richer boundaries, but generic multivariate node functions may obscure how individual features affect routing. We introduce Compound Additive Sparse Trees (CAST), a differentiable single-tree model that forms each split by summing learned univariate feature responses. Node-specific combinations of a shared nonlinear basis and direct linear terms give these responses flexible shapes while preserving an exact feature-wise decomposition of the routing score. Stochastic gates sparsify the feature responses used at each node, and exact folding replaces constant subtrees with leaves. On the complete 20-dataset benchmark, CAST achieves the best mean rank (1.100) in the primary five-method comparison. It ranks first at each matched depth setting from three through six. Plots of selected learned nodes show their nonlinear split boundaries and individual feature contributions in original feature coordinates. Across the primary runs, exact folding removes 30.3% of the starting trees' internal nodes in aggregate while preserving the selected gated trees' predictions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.