acceptodds
Under review as a conference paper at ICLR 2027

Hierarchical Runtime Alignment for Safe Long-Context Agents

Abstract

Long-context agents require safety controls that remain effective as model state, retrieved context, and cache state change over time. We introduce hierarchical run- time alignment, a deployment-time framework that combines three coupled layers: L1 parametric safety adaptation, L2 contextual policy evidence, and L3 cache-level safety preservation. Using matched vanilla/abliterated pairs from Qwen3-8B and Llama-3.2-3B, we show that runtime safety is a state-dependent calibration prob- lem rather than a single guardrail problem. Our experiments yield three findings. First, contextual policy substantially reduces harmful compliance in the tested abliterated models. Its false-refusal cost varies with model state and policy strength. Second, the tested safety–utility operating point depends on model state. Third, matched benign controls distinguish cache-allocation effects from uniform preci- sion reduction. Generic INT4 changes harmful and benign behavior together. The Qwen allocation control changes which signals retain precision. Key comparisons use three seeds, paired uncertainty, two independent semantic judges, and a prevali- dated fixed-subset LongSafetyBench evaluation. Across factorial sweeps, agent tasks, adaptive attacks, many-shot jailbreaks, RULER, and LongSafetyBench, the results support joint, model-state-aware calibration rather than a universal all-on configuration.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.