How Stable Are LLM Errors? Failure-Conditioned Decision Stability in LLMs
Abstract
Large language models produce fluent but incorrect answers, but wrong answers need not be equally stable under small changes in query context. We study this failure-conditioned decision stability on candidate-mappable closed-answer failures using a fixed weak re-query that provides no gold label, feedback, or additional evidence. For each failure, we compare the gold candidate with the realized wrong candidate and use their final-window candidate margin as the main pre-query coordinate; full-response traces are treated descriptively, and post-query quantities only as response diagnostics. Across four tasks, recovery in the two primary models rises from near zero in the most failure-favored final-window states to 39.4-60.2% in the largest-margin bins, with the same Q5-Q1 ordering in every task; the remaining models expose weaker, sparse, re-query-sensitive, or negative boundary cases. A replay-matched audit of 607 Qwen failures removes post-answer continuation from the readout: the literal answer-token margin yields 0.794 AUROC versus 0.800 for the original final-window coordinate. On the same 607 replay-matched Qwen failures, confidence-controlled analysis yields 0.789 AUROC for the margin versus 0.661 for wrong-candidate probability. A separate response decomposition further shows that the margin is associated both with whether an erroneous answer changes and, among changed non-binary answers, with whether the new answer is gold rather than another wrong candidate. We position residual-error stability as an offline analysis axis rather than a deployment-time error detector.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.