PriorBPB: Reducing the Target Length Confound of Bits-per-Byte While Improving Accuracy Prediction
Abstract
Bits-per-byte (BPB) normalises log-likelihood (in bits) by bytes rather than tokens to stay tokeniser-independent, and is widely used to compare language models across heterogeneous tasks. It also depends on how long the answer strings are. The bits a model spends on a gold string can be written as , with a fixed overhead, a per-byte rate, the least-squares fit in length, and the rest. Since the residuals sum to zero, total bits over total bytes is exactly , with the task's mean gold length. Because costs that do not scale with length are spread over the answer, tasks with long answers score better, as do long items within a task. Averaging per-item BPB has the same dependence. Across a 19-task suite on Qwen3 models, BPB correlates with mean gold length at Spearman to . OLMo-3-7B behaves the same way (Spearman ). Among the Qwen3 models we tested, the smallest model appears to score best (an apparent inverse scaling behaviour) due primarily to differences in the overhead. We propose a less length-biased metric, PriorBPB, which estimates a per-(model,task) marginal per-byte cost curve from per-token surprisal and averages it under one shared length prior instead of under each task's own length distribution. This cuts the mean absolute Spearman correlation with answer length from to across 60 base checkpoints from five model families, while correlating with accuracy as well or better on base and instruct checkpoints (mean Spearman ).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.