acceptodds
Under review as a conference paper at ICLR 2027

Conditional Hypothesis Generation for LLM-Based Text Analysis with Researcher-Specified Covariates

Abstract

An important goal of text analysis is to discover interpretable hypotheses about how language varies with outcomes of interest. Recent LLM-based methods select globally discriminative patterns that do not necessarily answer a researcher’s question. We introduce , the task of generating hypotheses that distinguish outcome groups within researcher-specified covariate strata. Covariates encode domain knowledge about which comparisons matter and can guide hypothesis discovery. We study two statistical challenges that arise in conditional discovery: relevant strata may be underrepresented within outcome groups (), and differences may reverse direction across strata (). Using sparse autoencoder features, we examine stratified regression, interaction models, and demeaning for covariate-aware selection. Controlled synthetic experiments show that these approaches can improve recovery of targeted differences overlooked by global selection, with effectiveness depending on the structure of conditional associations. On two real-world datasets, experts rate the covariate-aware method’s unique hypotheses as more helpful on average than the global baseline’s. We also illustrate how conditional comparisons can help researchers interpret generated hypotheses and formulate new ones.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.