Nuisance-Induced Confounding in LLM Output Sensitivity Analysis: A Semantic Normalization Framework
Abstract
Distribution-based sensitivity analysis aims to quantify how large language model (LLM) outputs change under input perturbations, but raw output distributions can conflate semantic changes with nuisance variation from stochastic generation, surface form, language, and representation choice. We formalize this nuisance-confounding problem by modeling LLM outputs as mixtures of meaning and nuisance factors, showing that raw-output distributional sensitivity can be nonzero even when the underlying meaning is unchanged. Motivated by a nuisance-invariance criterion for meaning-centered sensitivity analysis, we formulate a semantic normalization framework that maps raw outputs into meaning-focused representations before distributional comparison. From a conditional information bottleneck perspective, this transformation is interpreted as approximately preserving meaning while suppressing nuisance-dependent variation. Experiments on controlled perturbation analyses show that meaning-preserving nuisance changes induce nonzero raw-output sensitivity, whereas summary-based normalization reduces this nuisance-induced sensitivity by 53.6% on average while retaining task-relevant content. In multilingual patient-data experiments with controlled semantic changes, summary-based normalization yields more comparable sensitivity estimates across languages. These results characterize nuisance confounding as a measurement problem and support summary-based semantic normalization as a practical mechanism for mitigating its effects.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.