acceptodds
Under review as a conference paper at ICLR 2027

SELECTIVE INVARIANCE UNDER STRATEGIC INFORMATION MANIPULATION

Abstract

An agent that buys on a user's behalf reads a seller's description, contract terms and framing, then accepts the offer or walks away. Such a rule must not move when the seller merely re-presents a fixed transaction, yet must still move when the transaction genuinely changes. The latent state is unavailable at inference, so the rule cannot tell the two apart by inspection; it must be built so that one class of change cannot reach it. We call this selective invariance and ask when it is attainable. Given a specified admissible channel and observable representation, comparing the displacements that channel can produce with the direction that carries utility gives a criterion that sorts channels into three classes: invariance is free, invariance costs enumerating a completion set, or invariance is unattainable and the channel can only be estimated. Existing structural defenses show how to isolate particular flows; our contribution is an ex ante test of whether exact model-independent closure exists for a specified channel, not a detector that discovers arbitrary free-text manipulations. Because an LLM buyer is an opaque model rather than one we design, a guarantee valid for every decision function requires exact invariance, which no prompt-level defense provides. We measure both halves on paired renderings of the same hidden transaction, across three second-hand marketplace and bargaining benchmarks and seven model families. Presentation alone raises harmful acceptance by even for a buyer that accepts no bad offer under the honest rendering, and a safety prompt's protection varies seventeen-fold across families. Structural closure cuts harmful acceptance from to at essentially unchanged valid acceptance, holds the closed channels at exactly zero, and still reproduces a hidden-state oracle's decision under genuine concessions on all held-out negotiations. Applied unchanged to tool-using agents inside AgentDojo, with their suite, their injections and their verdicts, the criterion predicts in advance which defenses can be exact: ours takes attacker success from to exactly on the channel it closes, at no cost in task utility, while the published datamarking baseline reaches .

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.