acceptodds
Under review as a conference paper at ICLR 2027

Beyond Accuracy: Stable Predictions, Unstable Topology in Transformer Representations

Abstract

Predictive accuracy remains the dominant metric for evaluating robustness in modern natural language processing (NLP) systems, often suggesting strong resilience even under adversarial or corrupted inputs. We argue that this view can obscure failure modes that emerge in the induced local topological structure of learned representations. Focusing on BERT-based classifiers, we study the stability of local structure induced by k-nearest neighbors (kNN) under semantically-preserving transformations, including adversarial obfuscations. Despite near-perfect classification performance across all conditions, we find that kNN neighborhood structure shows substantial changes, where local relationships between embeddings are rewired even when predictions and semantic content remain unchanged. We further introduce a normalization-based sensitivity measure that relates topological change to embedding perturbation magnitude and show that neighborhood topology changes disproportionately relative to embedding displacement across all conditions, with cosine displacement remaining near-zero while Rank Stability increases monotonically with perturbation strength. This effect is consistent across neighborhood scales and occurs without loss of label consistency, indicating that predictive decisions remain stable while representation structure varies. We quantify these effects using topology-based stability metrics and observe large, statistically significant differences between clean and corrupted inputs. Overall, we find a systematic mismatch between predictive performance and the stability of the induced neighborhood topology, suggesting that accuracy alone is insufficient to characterize robustness and that topology-based measures provide a complementary diagnostic for hidden vulnerabilities.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.