Towards Robust Reasoning Trajectories: Certifying Chain-of-Thought Prefix Stability under Local Perturbations
Abstract
Chain-of-thought (CoT) reasoning enables large language models (LLMs) to solve complex tasks through intermediate reasoning steps, but extended CoT generation also makes inference increasingly sensitive to local perturbations. Small changes in input prompts can alter the reasoning trajectory, making both its stability and computational behavior difficult to characterize. Existing work has largely studied this phenomenon empirically, leaving a fundamental question unresolved: how long can a CoT prefix be certified to remain unchanged under bounded perturbations? We address this question by formalizing trajectory stability as exact prefix preservation under bounded local representation perturbations. We derive finite-perturbation propagation bounds for causal Transformers and combine them with clean next-token decision margins to obtain sufficient conditions for preserving every token in a target reasoning prefix. The resulting certificate explicitly links perturbation location, perturbation magnitude, and certifiable prefix length, providing a trajectory-level robustness guarantee for autoregressive reasoning. We evaluate the theory on four reasoning models and three benchmarks, and observe stability patterns that consistently follow the predicted effects of perturbation location, magnitude, and prefix length. Together, our theoretical and empirical results provide a principled framework for certifying the robustness of reasoning trajectories in autoregressive language models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.