Beyond Token Entropy: Measuring Distinct 2 Forms Of Output Variation In LLM Reasoning
Abstract
Reinforcement-learning-based post-training has made output diversity an increasingly important concept in LLM reasoning, but “diversity” can refer to several behaviorally distinct forms of variation. We study three observable proxies: lexical richness, semantic dispersion, and equation-pattern variation. In an original three-model pilot across StrategyQA, MMLU College Mathematics, and ARC-Challenge, the proxies exhibit markedly different pooled associations with reasoning accuracy. A benchmark-centered reanalysis shows that the lexical-richness association remains comparatively stable (r ≈ 0.784), semantic dispersion is near zero (r ≈ 0.019), and the initially strong negative equation-pattern association is substantially attenuated (r ≈ −0.405 versus pooled r = −0.834). We therefore interpret the study as evidence that output diversity is measurement- and benchmark15 dependent rather than as evidence that any diversity axis causally improves or harms reasoning. We additionally report an expanded five-model evaluation spanning 1.5B and 7B checkpoints on GSM8K, MMLU College Mathematics, and ARC-Challenge. The expanded study exposes substantial variation in accuracy, generation truncation, and answer-extraction reliability, reinforcing the need to separate behavioral measurements from inference-pipeline artifacts. Together, these results motivate more careful, axis-specific measurement and controlled training interventions before diversity-based prescriptions are made for RL algorithms.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.