ALERT: Attention-based Long-context Early-warning via Routing Trajectories
Abstract
Long conversations can retain relevant evidence while a language model answers incorrectly. We introduce ALERT, which predicts retrieval failure before the first answer token. A learned ranker selects one history message using attention at the final prompt position; a second probe predicts failure from that source's attention profile. The LLM remains frozen, and warnings trigger re-injection of selected evidence while preserving the full history. Paired evaluations through 32K separate localization, failure prediction, and recovery across NoLiMa-style tasks, BABILong, turn-based RULER, and Scatter-MCQ. In an initial confirmation on 576 new BABILong stories at 16K, frozen monitors achieve failure AUROC 0.710 for Gemma-4-E4B and 0.759 for Qwen3-4B, versus 0.518 and 0.517 for raw next-token margin. For these two models, respectively, frozen-monitor AUROCs are 0.941 and 0.802 on RULER, and 0.672 and 0.662 on Scatter-MCQ. Selective re-injection improves accuracy by 14.9 and 17.0 percentage points on BABILong, and 9.5 and 5.7 points on Scatter-MCQ, respectively. These results support pre-generation failure prediction and net recovery, although neither BABILong policy exceeds always-re-injection accuracy. Source-selection errors and frequent warnings remain constraints on effective selective intervention.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.