Comparing Internal Signals for Harmful Prompt Detection
Abstract
Recent methods detect harmful prompts using the gradients or hidden representations of Large Language Models (LLMs). These signals capture different parts of how the model processes an input, and it is not clear how much each of them helps detection. We compare seven groups of statistical features computed from losses, gradients, logits, attention weights, hidden states, neuron activations, and normalization outputs. On ToxicChat and SafetyPromptCollections, we examine the feature distributions, the effect of removing each group, feature selection, and the choice of LLM. Removing the hidden state statistics and neuron activation rates lowers the average F1 the most, followed by removing the input embedding gradients. The effect of the other groups varies across datasets. We provide HELM, an implementation that combines these features with existing classifiers. The appendix gives the full feature definitions and further results.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.