Better Prediction Can Make a Worse Alarm: Auditing Predictive Selection in LLM-Agent Safety Monitoring
Abstract
Runtime monitors support the safe deployment of LLM agents by warning of impending failures under strict false-alarm budgets. Candidate monitors are commonly benchmarked and selected using offline predictive metrics, but deployment evaluates them through sequential warning policies that turn predictions into alarms. This practice assumes that a monitor that predicts failures better offline will also make the better runtime alarm. However, whether that assumption holds remains largely unexamined. We formalize this question as predictive-to-policy ranking transfer, which measures whether predictive preferences between monitors are preserved under a specified warning rule and false-alarm budget. We audit ranking transfer for five monitors across three warning rules and four false-alarm budgets, with extensions across candidate families, agent models and environments. We find that a consistently better offline predictor can become the worse alarm, even when both rankings are highly stable. These reversals can be costly: in one low-false-alarm setting, stable reversals incur a median 30.4% relative loss in held-out recall. In contrast, selecting monitors by warning performance improves held-out recall by 2.2–5.8 percentage points while reducing achieved false-positive rates. We further show that this mismatch is structural. Standard predictive metrics summarize prediction quality across observations, while warning performance also depends on how scores behave over an entire execution. Reordering scores within an execution can preserve pointwise predictive summaries while substantially altering warning performance under rules that depend on temporal order. Thus, predictive metrics alone cannot establish which monitor makes the better alarm; ranking-transfer analysis makes the prediction-to-decision handoff explicit and tests whether offline comparative conclusions survive sequential realization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.