INFER: Feature Engineering as Inference with Transferable Reinforcement Learning Agents
Abstract
Feature engineering, the transformation of raw columns into predictive features, is one of the most time-consuming steps of tabular machine learning. Automated feature engineering (AutoFE) faces an exponentially growing candidate space and expensive per-candidate evaluation, and nearly every existing method pays for a fresh search on every new dataset. We propose INFER (INductive Feature Engiמeering via Reinforcement), which learns feature engineering once, offline, and applies it to any new dataset by inference. INFER is a hierarchical deep reinforcement learning framework with two cooperating agents: an outer agent reads dataset meta-features and selects a small subset of promising columns, and an inner agent composes operators over that subset to build a new feature. Both agents are trained on a collection of datasets and then applied, with frozen weights and no retraining, to previously unseen datasets. Feature generation on a new dataset therefore costs only forward passes and a lightweight selection step, with the number of generation runs as an explicit compute budget. On 75 binary-classification datasets, under a strict nested protocol in which every candidate is generated and fitted on training rows only, INFER improves over the original features on 81–83% of datasets with both XGBoost and RandomForest. It outperforms the RL-based method DIFER under both classifiers and the state-of-the-art OpenFE under XGBoost, and matches OpenFE under RandomForest. Code: https://github.com/infer2026/INFER
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.