Blind Fibers: Safety-Sufficient Representations for Reliable Agent Verification
Abstract
Reliable evaluation of AI agents depends not only on which trajectories are tested, but also on what information about those trajectories is shown to the verifier. We study failures caused by representations that make safe and unsafe behaviors appear identical. We call such cases blind fibers and show that no downstream verifier can reliably distinguish them once the relevant safety information has been discarded. We introduce the Verifier Sufficiency Defect, which measures the safety information lost by a verifier representation, and establish its connection to irreducible verification error. Motivated by this analysis, we propose FIBER, a method that identifies outcome-matched safe and unsafe trajectories and learns a compact safety representation that preserves the distinctions missed by the original verifier view. Controlled tool-agent experiments validate the theoretical predictions and show that the repaired representation generalizes to unseen fault combinations. We additionally identify the same failure mechanism in a publicly released AgentDojo trajectory.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.