Early Warning for First Failure in Multi-Round LLM Generation via Survival Analysis
Abstract
Large language models increasingly operate through sequential generation, including multi-turn interactions with users and multi-step chain-of-thought reasoning. To model response-level reliability, existing uncertainty-quantification methods provide useful measures. However, once an early error enters the context, subsequent outputs may remain internally consistent with the corrupted trajectory, causing consistency-based uncertainty estimates to appear deceptively low and obscuring where the trajectory first breaks. This motivates the central problem of anticipating when the first failure will occur. To tackle this problem, we formulate first-failure monitoring as a discrete-time survival problem, scoring each step before its output is released. Beyond the process formulation, we incorporate propagated historical instability to capture pre-failure dynamics, and apply a conformal calibration layer that bounds, with finite-sample validity, the probability that a first failure occurs without a warning. We further show that access to propagated historical instability cannot worsen the oracle warning-burden frontier by score-class nesting, and derive conditions under which this additional history yields a strict burden reduction. Empirically, we evaluate the framework across four HalluHard domains with two open LLMs and on PRM800K reasoning chains, comparing against raw uncertainty scores, alternative feature sets, and two survival models. Within our hazard model, adding propagated historical instability improves both AUROC and AUPR in all 12 settings, and the calibrated propagation-aware rule achieves the lowest warning burden among admissible constructions across the evaluated operating points.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.