acceptodds
Under review as a conference paper at ICLR 2027

Comparing Internal Signals for Harmful Prompt Detection

Abstract

Recent methods detect harmful prompts using the gradients or hidden representations of Large Language Models (LLMs). These signals capture different parts of how the model processes an input, and it is not clear how much each of them helps detection. We compare seven groups of statistical features computed from losses, gradients, logits, attention weights, hidden states, neuron activations, and normalization outputs. On ToxicChat and SafetyPromptCollections, we examine the feature distributions, the effect of removing each group, feature selection, and the choice of LLM. Removing the hidden state statistics and neuron activation rates lowers the average F1 the most, followed by removing the input embedding gradients. The effect of the other groups varies across datasets. We provide HELM, an implementation that combines these features with existing classifiers. The appendix gives the full feature definitions and further results.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.