TrendGuard: Layer-Wise Safety Trend Learning for Pre-Generation Multimodal Guarding
Abstract
Multimodal jailbreaks can hide harmful intent in text, images, or their interactions, posing significant challenges for vision-language model safety. Existing safeguards either inspect generated responses, requiring full decoding before intervention, or monitor internal states before generation. While internal-state defenses are more efficient, existing approaches typically rely on isolated layer representations or order-agnostic aggregation, overlooking how safety evidence evolves across decoder depth. We present TrendGuard, a lightweight pre-generation guard for frozen large vision-language models (LVLMs) that explicitly models this evolution as a layer-wise safety trend. We observe that unsafe evidence follows distinct depth-wise patterns: layer-wise unsafe scores may rise, decline, or persist across decoder layers, and preserving such ordered trends enables discrimination between persistent unsafe evidence and transient risk-like activations. TrendGuard extracts last-token residual and MLP activations and learns sparse layer probes to obtain ordered unsafe scores across decoder depth. Based on these trends, we instantiate two representations: Trend-Abs, which captures absolute risk evolution through scores and interpretable trend statistics, and Trend-Rel, which further models class-conditional trend relationships using safe and unsafe prototypes. Experiments on seven LVLM backbones and four unsafe benchmarks demonstrate that TrendGuard achieves state-of-the-art low-FPR safety detection, improving AVG TPR@FPR≤0.05 by 6.36% on average over mainstream methods. Under the same false-positive budget, it achieves perfect recall on most benchmark–backbone pairs. Warning: this paper contains example data that may be offensive or harmful.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.