Guardrails Must Be Earned: Fabrication-Resistant Compilation of Failure Traces into Anti-Pattern Skills
Abstract
Agents that compile failures into guardrail skills install whatever the compiler writes, and self-generated skills lower performance by 1.3 points where human-curated ones raise it by 16.2. Existing admission rules credit a candidate for plausibility or a later success, not for preventing the failure it names. We propose Earned Guardrails (EG), which treats a self-authored guardrail as a causal claim to be tested and bounds the share of fabrications that the library accumulates over a prespecified horizon. Typed failure compilation states what a candidate claims in a typed failure record, whose frozen matcher decides when a replay counts as the same failure, pinning the claim under test before any evidence is seen. Counterfactual replay admission tests that the named failure recurs without the guardrail and clears with it, composing the two arms into a fixed-level e-value test whose level is priced in advance from the budget and the proposal horizon. A post-admission ledger then promotes or retracts each admitted guard on randomized-exposure evidence of its marginal benefit, keeping a guard installed until that benefit is certified negative. Every level and hyperparameter is fixed before the first admission decision, and no threshold is tuned on the evaluation streams. On AppWorld, Terminal-Bench, and a determinism-filtered SWE-bench-lite slice, EG holds the false guardrail rate (FGR) at 0.08 installed and 0.09 ever-admitted under and retains 0.92 of genuine guards, whereas suppression-only admission reaches an FGR of 0.25 at 0.92 retention. Across 30 independently randomized streams carrying 300 injected fabrications, EG admits zero of them, while every baseline rule admits phantoms in all 30 streams, the LLM judge included.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.