acceptodds
Under review as a conference paper at ICLR 2027

Blind Fibers: Safety-Sufficient Representations for Reliable Agent Verification

Abstract

Reliable evaluation of AI agents depends not only on which trajectories are tested, but also on what information about those trajectories is shown to the verifier. We study failures caused by representations that make safe and unsafe behaviors appear identical. We call such cases blind fibers and show that no downstream verifier can reliably distinguish them once the relevant safety information has been discarded. We introduce the Verifier Sufficiency Defect, which measures the safety information lost by a verifier representation, and establish its connection to irreducible verification error. Motivated by this analysis, we propose FIBER, a method that identifies outcome-matched safe and unsafe trajectories and learns a compact safety representation that preserves the distinctions missed by the original verifier view. Controlled tool-agent experiments validate the theoretical predictions and show that the repaired representation generalizes to unseen fault combinations. We additionally identify the same failure mechanism in a publicly released AgentDojo trajectory.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.