acceptodds
Under review as a conference paper at ICLR 2027

LinAlg-Bench: When LLMs Stop Computing — Structured Hallucination, Sign Drift, and Depth-Gated Collapse in Matrix Arithmetic

Abstract

Frontier LLMs execute linear algebra flawlessly — until, at a measurable depth, they stop computing and start fabricating. We introduce LinAlg-Bench, a diagnostic benchmark of nearly 1,200 SymPy-verified problems across nine task types and 3×3–5×5 matrices, evaluated on ten frontier models at temperature zero (over 11,000 graded outputs, with results stable across three complete runs). Failure is not gradual but structural: models near ceiling on shallow tasks collapse to near zero on 5×5 eigenvalues — and the collapse is spectrum-conditional, with over 40% aggregate accuracy on integer spectra versus under 15% on irrational ones, implicating required numerical depth rather than matrix size. A three-stage forensic pipeline reveals that failures dissociate cleanly by task: eigenvalues fail by constraint-aware confabulation — fabricated values still match the matrix trace in nearly half of cases; determinants fail by sign accumulation, with a dimensiongated complete collapse that is absent at 4×4 yet dominant at 5×5. Forcing an O(n3) algorithm yields no recovery, chain-of-thought prompting rescues only the one model already strongest on the task, and chained shallow operations with total arithmetic matched to one eigenvalue computation remain near-ceiling — together implicating depth-of-state, not operation count, strategy, or prompting, as the bottleneck. LinAlg-Bench frames this working-memory limit as a falsifiable hypothesis and provides the instrument to test it.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.