Rethinking Optimization Methods for Hierarchical Feature Selection
Abstract
Rare-event forecasting often depends on interpretable conjunctions: a single keyword, location, or temporal cue may be weak, while their joint occurrence can identify a meaningful precursor. Hierarchical Tensor Network Lasso (HTNL) selects Numerical Conjunctive Features (NCFs) through a structured sparsity objective. However, the resulting optimization problem is challenging, making optimization design critical to its practical performance. We give a unified diagnosis of four HTNL optimization families: original majorization-minimization (MM), variational quadratic form (VQF), alternating direction method of multipliers (ADMM), and weighted group lasso (WGL). Our analysis distinguishes three key factors that jointly determine HTNL optimization behavior. First, VQF and MM are algebraically equivalent under matched variational updates and optimize a squared-penalty formulation. Second, ADMM optimizes the unsquared objective through consensus splitting; we replace its truncated FISTA proximal solver with a scalar bisection update (ADMM-Bisect) for more accurate treatment of near-threshold blocks. Third, WGL targets the same unsquared objective, while all methods face rapidly decaying active-set admission scores. Across two synthetic datasets and 11 civil-unrest and flu-surveillance tasks, VQF achieves the highest mean AUC with the most compact non-collapsing supports. ADMM-Bisect reduces measured runtime, while calibrated WGL avoids the empty-support failures of its original scale. Synthetic results further show how proximal accuracy, pruning, and search budget affect candidate recovery, especially for longer conjunctions. Code will be made publicly available upon publication.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.