NuRA: Numerical Manifolds in Pretrained LLMs Guide Test-Time Adaptation
Abstract
Despite advances in mathematical reasoning, reliable numerical prediction remains a challenge for large language models (LLMs). We find that pretrained numeric output representations exhibit angular relationships associated with magnitude and digit patterns, which we describe as the Numerical Manifold. The patterns extend across tokenizers and model families, including vision-language and math-specialized models. In this work, we propose Numerical Representation Adaptation (NuRA), which turns pretrained numerical geometry into a label-free learning signal for test-time adaptation. NuRA adapts hidden representations by minimizing angular dispersion, encouraging probability to concentrate on aligned numerical directions without labeled test data or additional offline training. We analyze how pretrained numerical relationships shape probability redistribution during adaptation. Experiments across numerical tasks and model families show broad accuracy gains with updates to approximately 0.01% of MLP parameters.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.