acceptodds
Under review as a conference paper at ICLR 2027

Taming the Edge of Chaos: The Dynamical Origin of Prompt Robustness in LLMs

Abstract

Large Language Models (LLMs) exhibit fragility to microscopic, semantic-preserving prompt perturbations. To uncover the physical origin of this vulnerability and its mitigation, we propose a paradigm shift: modeling the LLM residual stream as a discrete-time non-linear dynamical system. Within this framework, we mathematically unify a disparate ”zoo” of robust alignment strategies, ranging from Sharpness-Aware Minimization (SAM) and Muon to noise injection and Weighted Instruction Tuning (WIT). We prove that despite their distinct engineering heuristics, these optimizers converge on a singular physical mechanism: they act as dynamical dampers that drive the network's Quasi-Lyapunov Exponent into a strictly dissipative regime (). Furthermore, we establish the first causal link between continuous-time neural flows and the geometric phenomenon of Progressive Neural Collapse (NC). Rather than adopting the Unconstrained Feature Model (UFM) as an axiomatic assumption, we rigorously prove that localized dynamical contraction () physically ”unlocks” the UFM. This continuous contraction acts as a physical engine that exponentially dampens intra-class variance (), sculpting the terminal representations into a highly symmetric Simplex Equiangular Tight Frame (). During inference, these geometric vertices act as asymptotically stable attractor basins (). Extensive empirical evaluations confirm that this topological defense intrinsically absorbs adversarial lexical noise and structural jailbreaks, achieving state-of-the-art zero-shot prompt robustness without sacrificing foundational model expressivity.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.