acceptodds
Under review as a conference paper at ICLR 2027

Beyond Direct LLM Prediction: Open-World Reuse of Heterogeneous Tabular Learnwares

Abstract

Large language models (LLMs) can draw on broad knowledge for prediction. Specialized tabular models offer complementary value even as LLM capabilities improve: they encode statistical relationships often learned from private data, while their explicitly defined inputs and targets support statistical auditing and validation. However, heterogeneous feature requirements and incomplete user inputs complicate their reuse. The learnware paradigm makes independently trained models identifiable and reusable by equipping each with a statistical specification of its training distribution, without sharing the original data. Building on this foundation, for the first time, we enable heterogeneous tabular learnwares to address open-world prediction queries whose observed features need not match any model's input requirements. Key to this capability is our finding that local statistical relationships encoded in different specifications can be jointly exploited to infer missing inputs required by other models. Our framework aligns specifications along shared features into a query-dependent context, from which a pretrained tabular prior infers missing inputs for model retrieval and prediction. Model-specific completion uncertainty guides filtering before aggregation, while statistical support under the same context determines whether to return the ensemble prediction or abstain. Across 41 tasks, our framework outperforms direct and refined LLM prediction and LLM-assisted model reuse. Controlled comparisons further show that statistical relationships encoded in other models' specifications can improve missing-feature completion and downstream prediction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.