Agent-Designed Search Spaces for Tabular Machine Learning
Abstract
Despite the rapid progress of LLM-based agents for planning, code generation, and debugging, their practical value for tabular machine learning remains underexplored. In this paper, we investigate a concrete use case: whether state-of-the-art agentic AI systems can design extended HPO search spaces for established tabular models that outperform the standard published search spaces of these models. Specifically, we represent each tabular model as a modular pipeline covering preprocessing, embeddings, architecture, training, and inference. We then task the agent to propose candidate code implementations for each module and use a classical HPO algorithm to jointly optimize over these candidates and the model's default hyperparameters. Compared with the base HPO spaces, the expanded search spaces improve four of the five model families across a suite of 45 datasets, with average relative gains of 0.6%, rising to 2.0% on small-to-medium regression datasets. Notably, these gains are obtained at matched budgets: the enlarged spaces outperform the base under the same number of tuning trials and ensemble members. The gains transfer to the recent TabArena benchmark, where the tuned and ensembled agentic spaces improve the Elo of four of the five model families and the two strongest agentic ensembles surpass AutoGluon's ensemble of conventional models. Under matched budgets, neither an end-to-end agent loop nor a human-curated bank of AutoML-library components reproduces these gains, suggesting that the value of LLM agents lies in expanding the design space. The code, prompts, and generated search spaces will be released upon publication.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.