Sparse Attention-Head Patching for Cross-Notation Numerical Comparison: Donor-Position Transfer
Abstract
Prompt-final attention-head patches can restore mixed-notation answers by transferring a donor-selected option position. In stored Llama and Qwen confirmations, aligned donors improve held-out pair consistency from 0/16 to 14/16 and 1/16 to 15/16; summary-level swapped-donor counts are 61/64 and 30/32 presentations, respectively. A relabeled-option intervention shows that, in this Llama setup, target choices follow the donor-selected option position even when donor and target use different labels. This is an output-level transfer result in one model and prompt protocol, not evidence for a portable latent router. A separately preregistered, outcome-informed Llama follow-up finds that knockout of the frozen six-head set reduces signed margins beyond all 30 matched nulls and that operand-span K- and V-vector donor patches shift margins toward a counterfactual answer. The aggregate knockout loss is 0.880 (95% pair-bootstrap CI [0.805,0.962]), but is concentrated in digits/digits prompts (2.390 vs. 0.127 and 0.123 in reciprocal mixed-notation conditions). A later held-out screen across scientific, comma, zero-padded, and word formats met its 80% intact-accuracy floor in only 2/8 format-template cells and stopped before causal tests. A same-cohort chat-format sensitivity run failed donor competence (43/64), so transfer under the original double-BOS rendering remains untested. The evidence supports cross-label donor-position transfer in the standalone prompt setup and a bounded operand-linked effect in one Llama task, not a broad sparse numerical circuit. Unsupported older results are omitted.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.