acceptodds
Under review as a conference paper at ICLR 2027

Neural Barrier Functions as Guardrails for Adaptive Chain-of-Thought Steering

Abstract

Neural barrier functions (NBFs) provide a principled way to certify probabilistic safety and reachability in stochastic systems. Leveraging the inherent stochasticity of Large Language Model (LLM) generation, we introduce GuardSTR, a framework that formulates Chain-of-Thought (CoT) reasoning as a finite-horizon stochastic reachability problem, where successful reasoning corresponds to reaching latent states of correct final answers. Based on this formulation, we learn a reasoning barrier function as a guardrail to steer CoT reasoning. Theoretically, it provides probabilistic guarantees under a set of principled conditions for the desired terminal outcomes. Practically, it assigns a quantitative feedback credit for adaptive activation steering at every reasoning step. GuardSTR requires neither parameter fine-tuning nor ground-truth answers during inference, ensuring compatibility with different open-source LLMs. Across four LLMs and seven tasks from BBH and MMLU, GuardSTR improves final-answer accuracy over unsteered baselines on all 28 model–task pairs, yielding a macro-average gain of 5.91 percentage points. It achieves maximum gains of 12.8 and 11.6 percentage points over two state-of-the-art methods, Static CAA and K-CAST, respectively.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.