acceptodds
Under review as a conference paper at ICLR 2027

Disambiguating Contextual Integrity Parameters through Uncertainty Quantification

Abstract

Recent work has shown that ambiguity in Contextual Integrity (CI) parameters can substantially impair LLM-based privacy judgments. However, it remains unclear which CI parameter should be clarified and when clarification is worthwhile. We introduce the Contextual Integrity Clarification Framework (CICF), a framework that estimates parameter-level uncertainty by generating candidate specifications of individual CI parameters and measuring their effects on the model's privacy prediction. The framework uses these estimates to select parameters for clarification. We evaluate CICF on the CI benchmarks PrivacyLens+ and ConfAIde+ with three LLM-based privacy judgment models. Our experiments demonstrate that CICF can support more reliable LLM-based privacy judgments through targeted clarification: it consistently identifies more effective clarification targets than random or low-uncertainty baselines, leading to an accuracy improvement of up to 12.0 percentage points in PrivacyLens+ and 6.3 in ConfAIde+ under ideal error detection. We also evaluate a supervised probe for error detection, which combined with our framework still improves accuracy by up to 7.1 percentage points in PrivacyLens+ and 5.7 in ConfAIde+. Sequential adaptive target reselection further improves estimated oracle-gated accuracy beyond one round in all three PrivacyLens+ models, with model-dependent effects on ConfAIde+

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.