Efficient and Adaptable Detection of Malicious LLM Prompts via Bootstrap Aggregation
Abstract
Large Language Models (LLMs) have demonstrated remarkable capabilities in natural language understanding. However, they remain susceptible to malicious prompts that induce unsafe or policy-violating behavior. Existing defenses face fundamental limitations: black-box moderation APIs offer limited transparency and adapt poorly to evolving threats, while white-box approaches using large LLM judges impose prohibitive computational costs and require expensive retraining for new attacks. To address these challenges, we present BAGEL (Bootstrap AGgregated Ensemble Layer), a modular, lightweight, and incrementally updatable framework for malicious prompt detection. BAGEL employs a bootstrap aggregation and mixture-of-experts-inspired ensemble of fine-tuned models, each specialized on a different attack dataset. At inference, BAGEL aggregates predictions from a stochastically selected subset of ensemble members, optionally augmented by a lightweight learned router that predicts the most suitable member to include. When new attacks emerge, BAGEL updates incrementally by fine-tuning just a single small classifier (86M parameters) and adding it to the ensemble. We evaluate BAGEL on a held-out, source-disjoint test set constructed entirely from datasets never used for training or calibration, ensuring that reported performance reflects generalization to out-of-distribution attacks. In single-turn evaluation, BAGEL achieves an F1 score of 0.757 at a size equivalent to 430M parameters, improving to an F1 of 0.772 when augmented with a lightweight learned router using fewer ensemble members (324M parameters). These results outperform popular black-box moderation APIs and most white-box baselines at lower costs. In multi-turn evaluation, BAGEL achieves an F1 of 0.660, outperforming billion-parameter baselines. Our results show ensembles of small finetuned classifiers can match or exceed the performance of billion-parameter guardrails while offering the adaptability and efficiency required for production systems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.