Watch and Correct: Efficient Streaming Intervention for LLM Alignment with Distribution-Free Guarantees
Abstract
Advances in large language models (LLMs) are enabling increasingly capable interactive applications, placing safety and responsiveness at the center of deployment. Despite progress in alignment, these models can still produce harmful content, motivating inference-time safeguards that detect and correct unsafe generations. However, repeated correction consumes computation and delays delivery, with both its effectiveness and cost depending on the generation and guardrail models. The unresolved deployment challenge is to determine the safety attainable within a limited repair budget and the intervention cost required to meet a prescribed risk tolerance. We introduce Watch and Correct (WaC), a framework for risk-constrained generation that combines statistical calibration of repair budgets with efficient streaming correction. WaC identifies budgeted intervention policies that meet a target harm rate with finite-sample, distribution-free guarantees, and characterizes the corrective effort they incur across model configurations. This connects safety requirements to concrete choices of generation models, guardrail models, and repair budgets. To make these policies practical for interactive use, WaC integrates feedback-guided repair with asynchronous verification and computation reuse, reducing redundant work and interruptions to response delivery. Experiments demonstrate that WaC reduces token lag by approximately half compared with state-of-art intervention methods, improving the responsiveness of guarded generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.