acceptodds
Under review as a conference paper at ICLR 2027

Similar Accuracy, Different Mechanisms: Training History, Scale, and Task Formulation in Transformers for Alzheimer's Disease Classification

Abstract

Similar predictive performance can conceal different computations and clinical decision rules. We study how training history, scale, and task formulation shape decoder-only transformers for ADNI-derived cognitive-status classification. We compare pretrained-and-LoRA-adapted Pythia models from 14M to 1.4B parameters with from-scratch GPT-NeoX controls, including architecture-matched configurations. Both training regimes are evaluated using classifier-head and next-token formulations. Circuit extraction, probing, activation patching, and clinical counterfactuals distinguish represented information from evidence that influences predictions. The 12-layer from-scratch classifier reaches 90% circuit faithfulness using 7 ± 5% of its attention heads, compared with 70 ± 21% for the architecture-matched Pythia-160M classifier; their corresponding macro-F1 values are approximately 0.82 and 0.70. Both estimates use the same mean-ablation procedure and two-point normalisation, with only valid extractions included. Readout patching places peak attention-mediated influence later in the pretrained-and-adapted members of the matched pairs. Linear probes select the same best-encoded cognitive feature in four of five domains across all controls and seeds, yet interventions reveal greater prediction sensitivity to verbal Functional Activities Questionnaire (FAQ) descriptors than to numerical-score edits in most evaluated from-scratch runs, with a seed-dependent exception. Length-matched clinically neutral filler produces few flips in most runs, but up to 59% in the Pythia-1.4B generative formulation. Score-preserving wording changes flip up to 100% of eligible, initially correct cognitively normal predictions in several models, revealing wording sensitivity that persists in the larger pretrained transformers evaluated here. Leave-one-out necessity analyses identify recurring influential FAQ items, including remembering appointments, assembling tax records, and writing checks, under the tested intervention contexts. Models with similar aggregate performance therefore exhibit different internal intervention profiles while showing recurring reliance on functional descriptors.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.