ReasonOps: Operator Segmentation for LLM Reasoning Traces
Abstract
Chain-of-thought traces from large reasoning models can span tens of thousands of tokens, yet we lack a vocabulary for describing their internal structure. Previous methods developed to analyze chain-of-thought traces are either too rigid or not expressive enough, failing to capture features across domains and models. To remedy this, we introduce ReasonOps, an unsupervised method that discovers a shared vocabulary of discourse operators for segmenting and annotating chain-of-thought traces. Using ReasonOps, we analyze 44,662 traces from 12 thinking LLMs spanning 6 families across 8 reasoning benchmarks and discover that they share a common compositional structure: 7 recurring reasoning operators—discourse-level moves such as Backtracking, Inferring, and Hypothesizing—that emerge from unsupervised clustering of sentence-initial 3-token pivots. These operators appear across every model family and benchmark domain, confirmed by three independent LLM judges who classify held-out samples at 70–76% accuracy. We analyze the structure of operators on easy vs. hard problems, revealing that the association between reflective operators and correctness changes with problem difficulty. Reasoning traces are highly model-identifying: structural operator features plus anchor-phrase text features recover the source model with macro-AUC , revealing that each model family has a distinctive reasoning fingerprint. Structural operator features predict within-problem answer correctness well above baselines. Classifiers built on these operators reach WP-AUC globally and on AIME. ReasonOps further enables early quality estimation well before the trace completes: we predict at WP-AUC for only 50% of the trace. Comparing OLMo checkpoints with a common base shows distinct operator distributions across post-training recipes, with greater backtracking under math- and code-focused training. The ReasonOps pipeline is unsupervised and annotation-free, enabling deep insights into LLM reasoning traces as well as strong downstream results on model identification and correctness prediction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.