acceptodds
Under review as a conference paper at ICLR 2027

LLMs as Feature Engineers: Simplifying Agentic Relational Learning

Abstract

Relational forecasting tasks, such as predicting customer churn from a database of customers, transactions, and products, can be solved by transforming the database into a single table of entity features and training a standard tabular model on it. The main difficulty of this approach is feature engineering, and recent works delegate it to large language models (LLMs). In particular, RelAgent gives an LLM agent control over the entire pipeline: the agent explores the database, writes SQL feature programs, chooses the downstream model and its hyperparameters, and revises all these decisions using validation feedback over many iterations. In this work, we ask which of these responsibilities actually require an LLM. We study a simple pipeline in which the LLM acts only as a feature engineer. Given a static description of the database, it writes several SQL feature programs independently, using a single request per program and receiving no predictive feedback. Model tuning and aggregation of predictions are handled by standard methods. On classification and regression tasks from RelBench, this pipeline outperforms the features engineered by an expert data scientist on average and stays close to RelAgent ( vs. average AUROC on classification). At the same time, one of our programs requires fewer LLM requests and about fewer processed input tokens than one search rollout of RelAgent. When a top-tier proprietary LLM is replaced with a much smaller open-source one, the average AUROC of our pipeline decreases by points, compared with points for RelAgent. Our analysis shows that a static database description is as informative as interactive exploration on most tasks and that conventional tuning further improves frozen LLM-generated features. It also shows that validation scores are a weak guide for choosing between programs: on classification, averaging independent programs outperforms selecting one of them by validation score. These results suggest that LLMs are most useful in relational learning as feature engineers, while model tuning and selection can be left to conventional methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.