Hierarchical Runtime Alignment for Safe Long-Context Agents
Abstract
Long-context agents require safety controls that remain effective as model state, retrieved context, and cache state change over time. We introduce hierarchical run- time alignment, a deployment-time framework that combines three coupled layers: L1 parametric safety adaptation, L2 contextual policy evidence, and L3 cache-level safety preservation. Using matched vanilla/abliterated pairs from Qwen3-8B and Llama-3.2-3B, we show that runtime safety is a state-dependent calibration prob- lem rather than a single guardrail problem. Our experiments yield three findings. First, contextual policy substantially reduces harmful compliance in the tested abliterated models. Its false-refusal cost varies with model state and policy strength. Second, the tested safety–utility operating point depends on model state. Third, matched benign controls distinguish cache-allocation effects from uniform preci- sion reduction. Generic INT4 changes harmful and benign behavior together. The Qwen allocation control changes which signals retain precision. Key comparisons use three seeds, paired uncertainty, two independent semantic judges, and a prevali- dated fixed-subset LongSafetyBench evaluation. Across factorial sweeps, agent tasks, adaptive attacks, many-shot jailbreaks, RULER, and LongSafetyBench, the results support joint, model-state-aware calibration rather than a universal all-on configuration.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.