FAROS: Characterizing Coding-Agent Search for Executable Time-Series Anomaly Detectors
Abstract
Coding agents can construct executable anomaly detectors from labeled time-series data. However, their open-ended search can vary substantially across runs in both length and resulting detector quality, making systematic evaluation difficult. We introduce FAROS (Framework for Agentic Run Observation and Synthesis), an evaluation framework that records search states during ongoing coding-agent search and applies a detector synthesis procedure to obtain detector readouts from selected states. This enables comparison of detector quality at selected points during search. It also separates variability across independent search states from variability in repeated synthesis from a fixed state. Using FAROS, we study detector construction from few labeled anomaly events and how reasoning setting and search effort affect detector quality and repeatability. Our experiments yield three findings. First, coding agents can construct detectors competitive with specialized pretrained time-series anomaly detectors from few labeled anomaly events. Second, at larger common output-token targets, stronger reasoning achieves higher mean detection performance, while longer lower-effort search does not recover the same performance. Third, we observe contrasting changes in search-state and synthesis variability across reasoning settings and the two tested label budgets as search proceeds. At the largest output target, search-state variability remains the larger component across reasoning settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.