A Multiclass Decision Tree with Balanced Topology via Label-Constrained K-means Grouping and SVM Learning
Abstract
In multiclass scenarios, effectively performing classification tasks with decision tree models is of great importance in fields such as medical analysis, intrusion detection, and data mining. This task faces numerous challenges, including accurate classification on large datasets and in imbalanced scenarios. Existing decision tree methods have the following shortcomings: (1) the learned tree topology is not sufficiently compact and requires pruning to control tree depth; (2) node splitting is performed through multiple sequential steps and has not been incorporated into a unified mathematical framework; (3) tree node splitting is often measured by distance or splitting criteria such as Shannon entropy, information gain, and the Gini index, without fully considering the global structure and information of nodes. To address these issues, we propose SLtree, a multiclass decision tree that obtains a balanced topology via label-constrained K-means grouping and learns nonlinear decision boundaries using support vector machines. During decision tree construction, SLtree introduces the clustering principle of the K-means algorithm, fully leveraging global relationships among samples. Moreover, label constraints are imposed to constrain the tree topology to a structurally optimized balanced binary tree, thereby yielding a compact topology construction. Node splitting can be mathematically represented by a constrained formulation and does not require multiple steps. To probe the classification limits of SLtree, we introduce an improved particle swarm optimization algorithm to individually optimize the parameters of each splitting node, yielding an improved variant of the algorithm, ISLtree. Experimental evaluation on 41 datasets demonstrates that our SLtree algorithm achieves a leading ranking among the compared algorithms, while its improved variant, ISLtree, significantly outperforms other methods. Moreover, ISLtree exhibits outstanding performance on imbalanced datasets and datasets with high classification difficulty, and it has clear advantages over other methods on datasets with large sample sizes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.