acceptodds
Under review as a conference paper at ICLR 2027

Beyond Max-Ent Targets: Disentangling Entropy and Diversity in LLM Post-Training

Abstract

Entropy-regularized trajectory-balance objectives have been proposed for LLM post-training as an alternative to GRPO-like objectives. By matching a distribution over trajectories rather than concentrating only on the highest-reward ones, they are expected to preserve diverse solutions, expand problem reachability, and improve accuracy on difficult tasks. We test these expectations under binary verifiable rewards. To expose the available target-design space, we expand the classical Shannon entropy with -exponential targets in a common Tsallis family, used here as an analytical and experimental degree of freedom. We show that group-normalized binary rewards reduce every reward tilt to a single incorrect-to-correct reweighting profile . This reduction reveals a coupling between support and scale, compresses apparently different target constructions into a narrow region, and provides a direct criterion for choosing trajectory-balance parameters. In the experiments, we separate nominal target geometry from learned-policy behavior and evaluate the latter after quantifying procedural variation. Although our sweep changes the -profile substantially, its behavioral effects reveal a clear separation between nominal target geometry and learned-policy properties: reachability and accuracy mostly remain within measured variation, while the strongest intervention reduces both. A matched GRPO comparison shows that Max-Ent objectives need not retain higher token entropy, even as an LLM judge reproduces the previously reported semantic-diversity effect. We thus identify nominal Max-Ent structure, token-level entropy, and semantic solution diversity as empirically distinct axes, showing that theoretical target changes cannot substitute for behavioral validation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.