Affective Non-Interference in LLM Agents: A Behavioral Dissociation Between Authorization and Execution Scope
Abstract
An agent may respond empathetically without changing which actions a user is authorized to request. Testing this separation requires identifying what changed: the behavior, the policy-relevant input, or its evaluation. We examine Affective Non-Interference (ANI) through three complementary experimental archives and two external benchmark audits. In a six-model single-agent study, the corrected mean policy-error contrast is percentage points (95% interval ), while a policy-delivery control changes errors by points. In a single-/multi-agent archive, re-scoring the same saved behavior under an explicit conflict-review rule changes the single-agent affect contrast within injected-conflict requests from to points. Yet no forbidden attempts occur in 10,296 restricted-task records. In a five-family archive, a -point ALLOW contrast cannot identify an affective scope effect: the current renderer removes the extra request in the combined condition. External audits expose parameter-declaration inconsistencies but also zero requirement coverage on 51 unseen tasks. These findings motivate an operational audit of input comparability, protected endpoints, and evidential coverage, without equating non-detection with safety or re-scoring with changed behavior.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.