acceptodds
Under review as a conference paper at ICLR 2027

Interpretable Discriminative Text Representations via Agreement and Label Disentanglement

Abstract

Interpretable text representations should be both predictive and usable as named constructs that independent analysts can apply. Existing approaches typically satisfy only one of these goals: anonymous discriminative bases predict well, while topic models and concept-bottleneck pipelines attach names to features but often validate interpretability only post hoc. We propose LFD (LLM-assisted Feature Discovery), which operationalizes interpretability through two testable properties: conceptual clarity, requiring a feature definition to transfer across independent annotators with high Cohen's , and label disentanglement, requiring the feature's meaning to be distinct from the target label. LFD discovers candidate features from contrastive, outcome-opposed text pairs, screens them using cross-LLM agreement, and selects only those that improve held-out residual predictive accuracy. We show theoretically that a -based clarity screen bounds annotation-noise-induced generalization-gap inflation; at the substantial-agreement threshold , the inflation factor is at most . Across ten text-classification tasks, LFD yields compact, named feature bases that approach full-embedding performance on rule-articulable tasks while using one to two orders of magnitude fewer dimensions. Human audits confirm that LFD features are substantially clearer and more label-disentangled than those from a leading LLM concept-bottleneck baseline: on iSarcasm, human-human agreement is for LFD versus for TBM, and human raters score LFD features as more disentangled on all ten tasks. These results suggest that interpretability in text representations can be made falsifiable through cross-rater reproducibility and explicit separation from the prediction target.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.