Identification Limits and a Protected Correction for LLM-Derived Guidance
Abstract
Downstream statistical and machine learning methods increasingly use LLM judgments as external guidance, such as penalty weights, pairwise comparisons or prior scores, often treating the elicited signal as having a stable meaning. Yet when an LLM is shown task-matched evidence, its judgment may combine an evidence-free preference with an evidence-induced shift. We show that repeated querying can estimate this composite judgment arbitrarily precisely without identifying its decomposition. Whenever the queried evidence configurations cannot be combined to reproduce the empty, evidence-free configuration, multiple preference-shift decompositions remain observationally equivalent. This ambiguity persists under adaptive querying and cannot be resolved by increasing the query budget. We therefore develop Selective Correction, whose guarantees depend only on observable quantities. A data-only reference procedure identifies comparisons that could change the decision, an evidence gate restricts queries to those supported by task-matched evidence and the LLM supplies their direction, with its certainty allowed only to attenuate the correction. We prove that any reference ordering whose margin exceeds the sum of its local correction budgets is preserved for every possible set of LLM responses. Across seven datasets, Selective Correction improves over the data-only Reference in every setting while certifying more than 93% of reference orderings in every audited setting. Randomization controls further separate the gains due to LLM content from those due to selective targeting alone.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.