HYSGUARD: RESETTING SAFETY STATE, NOT GENERATION, FOR STREAMING LLM SAFEGUARDS
Abstract
Streaming safeguards are becoming a necessary component of large-language-model serving because unsafe behavior can emerge only after decoding has begun.Existing monitors improve intervention latency by classifying prefixes or accumulating hidden-state evidence over time, but this temporal memory creates a second problem: after the safety regime changes, evidence from the previous regime can remain active. We identify this failure mode as safety-state hysteresis. It delays unsafe-onset alarms after long benign histories and, symmetrically, delays release after the model returns to safe behavior. The key observation is that generation con-text and safety-decision state have different persistence requirements: the language model should retain its full causal context, while the safety monitor should selectively forget stale risk evidence. We introduce HYSGUARD, a lightweight monitor that reads cross-layer hidden states, maintains fast and slow safety memories, detects regime boundaries from their discrepancy, and refreshes only the monitor state while leaving the backbone and KV cache untouched. We formalize the boundary signal, give an O(r)-state online algorithm, and define SWITCHTRACE, an evaluation suite covering safe-to-unsafe, unsafe-to-safe, and repeated switches. Controlled mechanism tests and a complete cross-model evaluation show how to distinguish temporal adaptation from ordinary sequence classification, while preserving the deployment advantage of cached autoregressive decoding.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.