acceptodds
Under review as a conference paper at ICLR 2027

What Carries Mathematical Computation? Tracing Functional Changes in Language Models

Abstract

Modern language models exhibit strong mathematical capabilities, yet fine-tuning on specific tasks can compromise performance on other mathematical problems. Can the changes supporting acquired mathematical skills be isolated within a compact update? We propose SAE-guided update selection, which connects changes in internal representations to the selection of parameter updates. A shared sparse autoencoder (SAE) identifies feature changes between the base and fine-tuned models. We then score rank-one terms of the LoRA update by their effects on recovering fine-tuned feature values and predicting correct solutions, retaining a subset under a fixed storage budget. Empirically, across four model pairs from the Qwen, Gemma, and Llama families, incorporating feature matching improves selection over prediction loss alone. The joint criterion surpasses the full adapter on addition, subtraction, and variable isolation on three pairs, while also preserving more GSM8K accuracy. Beyond selection performance, we examine whether the identified features contribute to the computation itself. Removing them impairs later computation and final answers even after supplying the correct next token. Moreover, with the selected update applied, restoring features associated with addition carry to their base values selectively disrupts carry-dependent solutions. Together, these findings connect the retention of mathematical skills in compact updates to internal features that support successive computation steps.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.