acceptodds
Under review as a conference paper at ICLR 2027

TargetGuard: Freshness-Certified Target Integrity for State-Changing LLM Actions

Abstract

An API call can succeed while modifying the wrong entity. TargetGuard makes this distinction explicit through a target-integrity benchmark contract and a freshness-certified selective execution gate. Given a proposed action and a public typed relation, the gate admits a unique evidenced target only after revalidating its relation, pre-state, and executor-version certificate; otherwise it abstains. We establish conditional target integrity when the relation links candidates to the requested entity and the executor protects validation and mutation atomically. Observationally indistinguishable targets require abstention or additional evidence. Our primary comparison uses 192 controlled Git and SQLite cases. Raw dispatch records 96 wrong-target writes. Both TargetGuard and an independently implemented typed revalidation reference record none, retain all 96 uniquely identifiable operations, and abstain on all 96 ambiguous cases, without additional LLM calls. Both reject all 96 actions made stale before revalidation in a separate stratum. These tests measure stale-state detection, not protection against concurrent changes between validation and mutation. Forty-two application replay cases and an eight-case resolver boundary audit test adapter behavior. A controlled, oracle-assisted 12-pair intervention isolates target substitution under fixed model responses. A separate 1,800-case analytical suite compares policy decisions, not persisted writes. The result is a reusable contract for measuring correct-target execution and the valid work retained by execution policies. It does not establish an advantage over typed revalidation or estimate deployment failure prevalence.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.