CINFA: Contextual-Integrity Norm Factorization for LLM Agent Actions via Counterfactual Commit Supervision
Abstract
Contextual privacy in large language model (LLM) agents depends on the recipient, content, and conditions of a proposed disclosure. We present CINFA, an action-conditioned framework for assessing contextual privacy before external actions are executed. CINFA represents each candidate action as a context-action commit that combines available context with its bound recipient, payload, and transmission evidence. Matched counterfactual edits provide global and partially observed factor labels for risk heads trained on frozen LLM representations. The runtime integrates context-guided generation with action checking, factor-guided revision, and re-checking, withholding unresolved actions. Across CI-RL-Synth, PrivacyLens, and ConfAIde, CINFA’s reported average leakage rate is 5.4%, compared with 12.4% for CI-GRPO. Its adjusted leakage rate, measured on helpful outputs, is 7.8% versus 17.3%, while task utility is 2.6 percentage points higher. Component comparisons, controlled factor edits, metadata stress tests, and evaluations across five frozen backbones characterize the evaluated configurations and their runtime inputs. These comparisons assess the integrated system rather than isolate individual mechanisms, positioning action-conditioned checking as a complement to generation-time privacy alignment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.