Benchmarking Sample Reasoning Quality from Internal Geometric Dynamics in LLMs
Abstract
Understanding and benchmarking how well a large language model reasons over a given sample is crucial, yet it is rarely legible from the model's surface output. Prior work overlooks how reasoning evolves across the model's layers, which we argue is indispensable to judging reasoning quality. Intuitively, just as a Rubik's cube can be solved by many strategies, the outcome is only one aspect of reasoning; whether the underlying strategy is concise and direct matters just as much. From this view, we propose Cube, a framework that benchmarks a sample's reasoning quality directly from internal geometric dynamics. By tracking the reasoning trajectory in a low-dimensional reasoning subspace, Cube identifies the characteristic cross-layer transformations and summarizes them with geometric signatures. Cube has broad applications, and in this work we present Cube-Eval and Cube-Selector for demonstration: one assesses reasoning quality directly from the model's internal dynamics, without any ground-truth labels, and the other pinpoints samples of low reasoning quality for targeted training. Extensive experiments demonstrate the effectiveness of both Cube-Eval and Cube-Selector.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.