acceptodds
Under review as a conference paper at ICLR 2027

The Geometry of Word Usage: Training-Free Detection of LLM-Driven Social Bots

Abstract

The proliferation of large language models (LLMs) has enabled a new generation of highly human-like social bots, posing significant challenges for reliable detection. Social bot detection is a user-level problem from many fragmented and multi-topic user histories. However, existing AI-text detectors operate at the passage level and degrade under such conditions. We propose a fundamentally different perspective: instead of modeling text likelihood, we characterize the geometry of repeated contextual word usage across a user’s history. Building on this insight, we introduce CDetect, a training-free framework that captures user-level semantic concentration. Specifically, CDetect encodes historical posts, aggregates contextualized representations of content words into representative directions, and models their distribution via a mixture of von Mises–Fisher distributions on the unit hypersphere. The resulting mixture-weighted diversity score quantifies whether a user exhibits semantically dispersed (human-like) or concentrated (LLM-like) language usage patterns. On the theoretical side, we establish finite-sample guarantees for the estimated diversity score, explicitly characterizing how vocabulary coverage and repeated contextual occurrences govern its statistical reliability. Extensive experiments on BotSim-24 and two TwiLLM benchmarks demonstrate that CDetect consistently outperforms both training-based and training-free baselines, achieving an average AUROC of 0.9416 and AUPRC of 0.9203. Further analyses confirm its robustness across varying history lengths, text budgets, and modeling choices.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.