RelCraft: Credit-Guided and Point-in-Time-Safe Program Search for Relational Learning
Abstract
Large Language Model (LLM) agents can automate relational feature engineering by searching over executable SQL, but free-form generation often produces invalid or temporally unsafe programs, scalar validation feedback poorly identifies which feature block helped, and repeated reuse of one validation set causes optimistic selection. We introduce RelCraft, a reliability-oriented framework with three matched components. A schema-typed grammar and deterministic point-in-time verifier restrict search to executable, row-aligned programs. Paired cross-fitted coalition evaluation estimates each block's average marginal utility under matched model fits. A protected split then selects a frozen SQL–predictor pair using a fixed complexity prior without returning selection outcomes to the LLM. We give a finite-sample guarantee for the bounded group loss used by this selector and evaluate predictive quality, search behavior, explicit future-edge violations, and deployment cost on 32 tasks from three relational benchmark suites.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.