acceptodds
Under review as a conference paper at ICLR 2027

CINFA: Contextual-Integrity Norm Factorization for LLM Agent Actions via Counterfactual Commit Supervision

Abstract

Contextual privacy in large language model (LLM) agents depends on the recipient, content, and conditions of a proposed disclosure. We present CINFA, an action-conditioned framework for assessing contextual privacy before external actions are executed. CINFA represents each candidate action as a context-action commit that combines available context with its bound recipient, payload, and transmission evidence. Matched counterfactual edits provide global and partially observed factor labels for risk heads trained on frozen LLM representations. The runtime integrates context-guided generation with action checking, factor-guided revision, and re-checking, withholding unresolved actions. Across CI-RL-Synth, PrivacyLens, and ConfAIde, CINFA’s reported average leakage rate is 5.4%, compared with 12.4% for CI-GRPO. Its adjusted leakage rate, measured on helpful outputs, is 7.8% versus 17.3%, while task utility is 2.6 percentage points higher. Component comparisons, controlled factor edits, metadata stress tests, and evaluations across five frozen backbones characterize the evaluated configurations and their runtime inputs. These comparisons assess the integrated system rather than isolate individual mechanisms, positioning action-conditioned checking as a complement to generation-time privacy alignment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.