Building Geometry-Aware Probes for Fine-Grained LLM Safety Classification
Abstract
Internal activation probes offer an efficient approach to safety monitoring in large language models, but most existing work focuses on general harmfulness rather than fine-grained safety categories. We study how harmful categories are organized in representation space and how this geometry affects classification. Category directions relative to safe inputs exhibit a strong shared orientation while retaining category-specific variation. Categories that align more strongly with this shared geometry also tend to be harder to classify, with Spearman correlations of -0.691 for mean inter-category directional similarity and -0.682 for shared-axis alignment. Motivated by this observation, we introduce Spatial-Sign-Conditioned Angular Kernel (SSCAK) probing. SSCAK estimates shared within-class directional scatter using spatial signs, conditions the representation space accordingly, and learns nonlinear category boundaries on the unit hypersphere using an angular kernel. Across seven instruction-tuned models and three token-aggregation strategies, SSCAK consistently improves macro-F1 over the corresponding standard probes. These results connect the geometry of internal safety representations to fine-grained classification and demonstrate the value of explicitly modeling within-class directional structure.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.