MASI: MECHANISM-AWARE INTERVAL INFERENCE WITH LLM SEMANTIC COVARIATES
Abstract
Insufficient sample size is a primary obstacle to precise inference, especially when observations pair text covariates with scalar outcomes. Unlike structured covariates with prespecified coordinates and interpretable perturbations, text lacks a canonical low-dimensional representation, and routine edits need not preserve outcome-relevant semantics. Modern large language models can generate source-conditioned texts, creating a new possibility for addressing this difficulty. Yet distribution shift and within-source dependence can make naive use biased and overconfident. We introduce Model-assisted Semantic Inference (MaSI), which uses generated texts to construct a more precise confidence interval for the parameter of interest while controlling the inferential effects of LLM-induced distribution shift under the stated calibration conditions. Specifically, the first set is generated under a scheme that preserves source content while varying expression style; together with real labeled data, it is used to learn outcome-relevant content coordinates and complementary style coordinates. A separate evaluation set constructs a calibrated auxiliary statistic in those coordinates. Density-ratio calibration addresses distribution shift, real-label residual correction preserves the population target, and uncertainty is estimated across original sources. Under explicit representation, calibration, and covariance-control conditions, MaSI provides consistent variance estimation and asymptotically valid confidence intervals. Simulations and empirical-data experiments show shorter intervals with near-nominal coverage in the primary settings, establishing a principled route to more precise inference from scarce text–outcome pairs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.