SCEPTER: Separating Action and Certificate Failures in Tool Agents
Abstract
A tool agent can choose the correct action and still be rejected because its evidence certificate is invalid. A successful retry may therefore reflect evidence repair rather than improved action selection, a distinction that aggregate task outcomes alone cannot resolve. We introduce SCEPTER, a controlled evaluation framework that separates these mechanisms by freezing initial proposals and varying evidence access and write scope, then evaluates certificate construction during public task continuation. In controlled workflows, deterministic evidence search matches grammar-constrained completion on all 278 action-correct, certificate-invalid proposals with constructible repairs, without a model call. An offline ablation reveals different recovery mechanisms: quote-only correction recovers most Qwen failures but none of the Llama failures, whereas binding-only correction recovers all Llama failures and a smaller subset of Qwen failures within the evaluated proof-local populations. In public benign continuation, direct compilation matches the observed task utility of certificate replacement with 31.9–34.8% fewer total model tokens after surplus-annotation normalization; replacement gains over rejection are supported for Qwen but inconclusive for Llama. Under attack, higher observed secure utility coexists with additional phishing deliveries, and no task-cluster contrast survives joint correction. A recorded-trajectory audit traces one such delivery to an endpoint rule that admits a malicious link from tool output despite a valid recipient certificate. These findings support compiler-first construction for the evaluated interfaces, not general semantic repair: constructing evidence can restore admission without correcting the task plan or the policy's treatment of harmful content.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.