BEAST: Benchmarking Evolving Abuse Streams
Abstract
Online abusive language evolves continually, so classifiers deployed in real-world moderation systems must continually adapt to text that may no longer resemble what they were trained on. Testing whether adaptation methods are effective requires a benchmark that preserves real-world task constraints—naturally occurring content, a fixed prediction task, a continuous stream, and meaningful volume—while methodologically isolating the challenge of continuous linguistic evolution. No existing continual-learning benchmark for abuse detection combines these properties. Instead, they manufacture distribution shift by sequencing tasks, datasets, bias dimensions, or adversarial perturbation types, which tests adaptation to one specified form of change rather than to language that evolves on its own. We introduce Benchmarking Evolving Abuse Streams (BEAST), a domain-incremental benchmark that holds the task and moderation policy fixed while only the incoming language changes. BEAST draws millions of naturally occurring posts from a single 3.5-year source and orders them so that incoming text grows progressively less familiar relative to the initial training period. Models are scored on an interval of the stream before that interval becomes available for training. We measure and report predictive performance, adaptation cost, and forgetting: across encoder and decoder models, increasing model scale and periodic retraining both improve average performance yet leave substantial end-to-end degradation, while retrieval augmentation performs far worse despite substantially higher compute. BEAST exposes a persistent gap between current adaptation methods and the robustness that evolving abuse detection demands.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.