acceptodds
Under review as a conference paper at ICLR 2027

EvoFG: Decoupling Operator Evolution and Feature Binding for Efficient Feature Generation

Abstract

Combining large language models (LLMs) with evaluation feedback has emerged as an effective approach to feature generation for structured data. However, directly generating concrete features typically requires jointly searching over operator structures and input-feature bindings. Consequently, potentially useful structures may receive negative feedback due to mismatched bindings, obscuring the direction of subsequent evolution and limiting LLM utilization efficiency. Some methods also include dataset names and column semantics in their prompts, introducing potential data leakage risks. To address these challenges, we propose EvoFG, a feature generation paradigm that decouples operator structure evolution from feature binding optimization. EvoFG searches over reusable, parameterized operator families, allowing the LLM to generate code without specifying concrete column combinations. Instead, a program enumerates candidate bindings and identifies high-quality feature instances through coarse-to-fine evaluation. The best binding, peak gain, and supporting evidence from other high-performing bindings jointly provide feedback to guide subsequent structural mutations, mitigating the misjudgment of structures caused by binding mismatches and turning each proposal into an exploration of multiple concrete features. This design delegates concrete column selection to programmatic evaluation and does not rely on dataset-specific information, thereby minimizing potential data leakage risks associated with LLM use. The high-quality features generated by EvoFG significantly improve the performance of different downstream models. Evaluations across multiple datasets demonstrate that EvoFG outperforms existing AutoFE methods. With results averaged equally across datasets under the same LLM budget, EvoFG produces over 30 times as many positive-gain candidate features and approximately 2.5 times as many ultimately selected new features as direct LLM-based feature generation, while reducing token consumption per selected feature by approximately 50%. These results demonstrate that organizing search around operator families and optimizing bindings programmatically enables more effective LLM utilization, greater search efficiency, and higher token efficiency, providing an effective search paradigm for feedback-driven evolutionary feature generation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.