acceptodds
Under review as a conference paper at ICLR 2027

Effect Survival: A Prospective Falsification Protocol for Agent-Safety Evaluation

Abstract

Agent-safety papers increasingly promote behavioral effects discovered inside configurable evaluation pipelines: a history increases unsafe action, a policy reduces it, or a representation improves utility. Benchmark-validity research shows that pipelines and scaffolds matter, but a complementary question remains: does the **effect claim itself** survive when its measurement is prospectively hardened and repeated? We introduce the **Effect Survival Ladder (ESL)**, a six-rung protocol spanning discovery, claim/estimand lock, measurement hardening, held-out breadth, replication blocks, and mechanism discrimination. A machine-readable **Effect Survival Card** records each rung's timing, decision rule, estimate, missingness, artifacts, and final claim license. We demonstrate ESL on two tool-agent safety effects across two prospectively frozen studies totaling 1,728 planned episodes over 12 semantic task families and three model families. First, development experiments suggested a large post-denial history effect (7/8 boundary crossings versus 0/8 fresh controls). After removing action-opportunity and state-delivery confounds, separating proposal from effect, and clustering inference by semantic family, a 1,008-episode confirmation produced 0/144 denied-history reattempts (pooled B0-A0 = -2.08 percentage points, 95% CI [-4.17, 0.00], exact family sign-flip p=0.25). Second, that confirmation exposed a prespecified -15.28-point authorized-utility penalty from rule-only context. A separately frozen 720-episode five-arm follow-up attenuated the primary unshammed F1-F0 replication contrast to -4.17 points (95% CI [-9.03, -0.69]), below its predeclared -10-point practical gate; the design also contained length-matched sham controls, but downstream mechanism claims were not promoted after the first gate failed. ESL is not a new significance test or a replacement for benchmark validity. It is a reusable protocol for exposing what stronger measurements a behavioral claim actually survived before the field builds on it.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.