acceptodds
Under review as a conference paper at ICLR 2027

Discovering Interpretable Text Categories with Natural-Language Codebooks

Abstract

Text classification often begins before a useful category inventory exists, and the inventory must also say how later texts are assigned. We formulate interpretable category discovery around three criteria: mutual exclusivity, collective exhaustiveness within a stated scope, and reproducible application from written rules. LLM-assisted Category Discovery (LCD) learns a natural-language codebook from unlabeled texts: each category has a name, a definition, inclusion and exclusion rules, and positive and negative examples. An LLM proposes every split, merge, revision and addition, and each is kept only if it also passes explicit tests on the memberships that an LLM reader verifies on the discovery texts; entries that claim fewer than three texts are retired. The frozen codebook is applied to new texts by a reader that may return zero, one or several categories, so gaps and overlaps remain visible. On BANKING77, CLINC150 and HWU64, whose development-split labels chose LCD's configuration (the label-free comparators ran untuned) and whose test split also scored intermediate configurations, LCD has the highest claim-set ARI and mapped macro-F1 over ten seeds on every dataset against every method we ran: TnT-LLM (ARI +.056, +.293, +.126; macro-F1 +.093, +.306, +.121; Holm-adjusted exact p ≤ .006), its distilled classifiers, every automated LLooM variant, every clustering rule with a data-chosen number of clusters, and four label-informed references. In refits, removing the addition operator lowers pooled ARI by .293 and the split operator by .079, and no removal raises ARI or F1 significantly. Because its reader labels 128 texts per request, labeling costs 11.4–13.7 times less than TnT-LLM's LLM labeler, and total cost is lower at every corpus size on two datasets and from 301 texts on the third. With a weaker reader, LCD stays ahead in macro-F1 on all datasets and in claim-set ARI on two (one with uncovered texts pooled); two-reader agreement is similar for both methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.