acceptodds
Under review as a conference paper at ICLR 2027

SemNorm: Semantic Normalization for Reasoning Traces Evaluation

Abstract

As large reasoning models generate increasingly lengthy and seemingly coherent natural language outputs, the difficulty of evaluating them at the trace-level has increased. In tandem, the verification of lengthy traces no longer provide sufficiently granular information about specific step-by-step errors or premise validity. To address these limitations in natural language reasoning traces, Semantic Normalization (SemNorm) introduces two components: 1. Atomic extraction, to process reasoning traces into atomic claims-level statements; and 2. Directed graph construction, which builds a knowledge graph from the pairwise entailment between atomic statements. This framework is founded on two observations. First, lexical-based methods are inadequate for establishing information content in diverse natural language generations. Second, embedding-based retrieval methods establish undirected semantic similarity, which leaves causal and implicative statements indistinguishable. To verify the information content in the extracted graph, SemNorm is applied as a scoring metric for reasoning quality and validity, and benchmarked against embedding-based methods as a semantic equivalence measure in perturbation-based experiments. Thus, this work discusses methods for estimating trace-level metrics, and introduces a directed reasoning graph representation for computing graph-derived scores which are comparably monotonic and on par with existing NLI-based scorers and embedding-based scorers for reasoning trace lexical, quality and correctness indicators and as a retriever model for RAG.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.