acceptodds
Under review as a conference paper at ICLR 2027

The Arithmetic Error Bar: A First-Passage Law for the Numerical Fragility of Language Models

Abstract

A benchmark score carries an uncertainty nobody reports: change the arithmetic that produced it — a numerical precision, a machine — and the number moves. On a 1000-question set the aggregate accuracy is unchanged, 66.2% against 66.1%, while 53 answers differ and 29 items change verdict; the cancellation that keeps the total still is an accident, and on a second machine of a different architecture the same kind of flips cost a full point. We call the resulting band the arithmetic error bar — here ±1.06 points — and we show that it is predictable in advance, item by item, from a single run, and that it can be paid down at a stated compute price. The mechanism is not the one usually assumed. Chaos predicts that reducing a pertur- bation makes the divergence arrive later; measurement shows that it removes it, and a state perturbation decays by an order of magnitude within ten tokens and never grows. Generation is a hybrid system — a contracting flow interrupted by a quantisation — and all of its sensitivity sits in the quantisation. The distance of a decision from its boundary is the geometric margin ρj = (ℓ1−ℓj)/∥w1−wj∥; normalised by the state norm it defines a threshold landscape along a generation, and the probability that the text diverges is a first-passage problem whose only parameter is the amplitude of the numerical noise, measured rather than fitted. It predicts which answers change (AUROC 0.960 against 0.797 for entropy, on held- out questions), places all 29 observed correctness flips in the half of the items it had flagged, and predicts 43 changed answers against the 53 observed; car- ried unchanged to the second machine it predicts exact agreement in float32, confirmed on all 1500 questions, and 71 divergences in bfloat16 against 66 observed. Run on two further models the law ranks well on one and not on the other. Rather than record this as a limitation we bound it: if the law holds, many questions are genuine coin tosses that no score can separate, so achievable rank- ing accuracy has a ceiling computable from the landscape and the amplitude alone, before any comparison of precisions is run. It depends on essentially one number, how often the arithmetic decides the answer, and against it all three models land at or near the best any score could achieve. Finally, re-running in float32 only the tenth of a corpus the law flags recovers 72.4% of the correctness flips and halves the error bar, with the residual predicted at every budget.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.