Monitor Bycatch: LLM Monitors Catch More than Intended
Abstract
As LLM agents become increasingly capable of enabling catastrophes, we need monitors that target rare, high-severity harms. Monitors can appear strong against benign content, achieving high recall with a <1% false positive rate, but are found to overflag when in deployment when they encounter other kinds of harmful content. We call these false positives . At the strict false-positive budgets of production systems, most false positives are bycatch: flags that are often harmful, but outside of their intended scope. We measure monitor bycatch on a new in-the-wild dataset of >240k publicly shared conversations that carries the natural incidence of each harm. We evaluate eleven monitors, each instructed to flag only its target harm category, using frontier models from several providers. At the observed in-the-wild occurrences of harm, precision falls below 50% in every target category and wrong-category fires outnumber true positives. Models can distinguish between harm categories, yet do not reliably respect those distinctions when monitoring for a single harm. On in-the-wild traffic, scoring every category of our harm taxonomy in one call reduces false positives from other harm categories by 10.3× compared to asking only about the target harm, at the same 95% recall. Finally, fine-tuning two open-weight models on the labels each produced when shown the full taxonomy, with the same model as a teacher, reduced bycatch even without the taxonomy at inference time.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.