Interpretable Feature Discovery from Unstructured Data Using LLMs and Statistical Choice Models
Abstract
Feature extraction from unstructured data typically relies on opaque embeddings unsuited to interpretable decision-making. Interpretable features, on the other hand, typically require costly, and often infeasible, hand-engineering. Large language models (LLMs) offer a natural bridge, as they can read raw data and describe it in interpretable terms. One could simply ask an LLM to extract features from a large sample of the dataset, but context limitations and diluted attention can lead to unsatisfactory results. We introduce a method that automatically discovers interpretable, decision-relevant features from unstructured data. Our key idea is to perform pairwise comparisons between items, asking an LLM to articulate, in interpretable terms, the features distinguishing each pair. We treat these explanations, or *rationales*, as data: we develop a statistical choice model that combines these rationales across the comparisons to identify relevant features and estimate per-item, per-feature scores. We first test our method on a synthetic hiring dataset of cover letters with known ground-truth features and show that it significantly outperforms direct LLM feature extraction in feature recovery. We then apply it to the ChangeMyView (Reddit) debate dataset, where it discovers several interpretable features associated with persuasion while reaching predictive accuracy comparable to established baselines. We also show how the same rationales can be used to audit LLM decisions, yielding a holistic score whose decomposition over interpretable features reveals what the LLM rewards.
est. 52% chance this paper gets accepted at ICLR 2027.
Recent trades
| Ago | Side | Shares | Outcome | Pricea |
|---|---|---|---|---|
| 2m | buy | 206.31 | Accept | 45.6% → 51.5% |
| 3m | buy | 2.19 | Accept | 45.5% → 45.6% |
a The traded outcome’s price before and after the fill.
All positions stay anonymous.