acceptodds
Under review as a conference paper at ICLR 2027

Cross-Entropy Loss Falls Short: What Cross-Entropy Loss Reveals about Large-Language-Model Performance?

Abstract

Scaling laws have made cross-entropy loss a central quantity for understanding and predicting the behavior of large language models, often with an implicit assumption that lower loss indicates better downstream performance. In this work, we revisit the reliability of this assumption on multiple-choice question answering (MCQA). Across diverse model families and datasets, we find that lower loss does not consistently correspond to higher accuracy. A key reason is that loss and accuracy depend on different quantities: cross-entropy loss measures how well a response is modeled, whereas accuracy is determined by whether the correct response has lower loss than its alternatives. We further observe that differences in loss-landscape geometry may make absolute loss values difficult to compare across model architectures and families. We then extend our analysis to an unsupervised setting, using the cross-entropy loss of the question context without ground-truth answers, and observe an even less consistent relationship with model performance. Motivated by the role of relative separation, we also examine loss dispersion across answer choices, but find that greater dispersion does not consistently indicate better performance. Together, our empirical and theoretical analyses reveal when and why loss-based quantities can become unreliable proxies for downstream model performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.