Relative Capability Organizes the Geometry of Answer Distributions in Test-Time Scaling
Abstract
Test-time scaling naturally produces distributions over candidate answers, revealing how a model's responses vary across repeated samples on the same problem. While prior work has connected individual properties of these distributions to difficulty, uncertainty, and aggregation performance, whether the distribution as a whole exhibits a shared structure across models and tasks remains unclear. We find that ranked answer distributions exhibit a shared geometry beyond ranking and probability constraints, systematically organized by model-problem relative capability. As relative capability increases, distributions progress from fragmentation, through structured competition among leading answers, to single-cluster dominance. Correct-answer location is reorganized within this geometry: it shifts sharply toward the leading cluster in the intermediate regime, after which answer dominance increasingly coincides with correctness. This empirical law generalizes to held-out models and tasks and extends to constrained answer spaces through a systematic geometric deformation. Together, these findings offer a new structural perspective on test-time scaling, revealing systematic organization in LLM reasoning behavior across models and problems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.