Better Aligned, Trains Worse: Intervening on Gradient Alignment in Local Learning
Abstract
Local learning rules replace backpropagation's transposed weights with a feedback matrix and are routinely checked by per-layer gradient alignment. A better-aligned matrix is read as one whose credit trains better, a premise tested so far mostly by correlation across runs. We show that alignment can prefer the feedback matrix that trains worse, and that following its preference lowers accuracy. We study the weight mirror, which learns its matrix from random probes. We run it at the published mirror learning rate with few probes and at a tuned setting with a larger rate and more probes. Each run also carries the other setting's matrix without using it, so both matrices are measured on the same forward weights. Measured this way, in a residual network with and without shortcuts, alignment prefers the matrix of the setting that trains worse. We then fork tuned runs. From one snapshot, one branch keeps its matrix and the other lets the preferred matrix compute the weight updates. That branch ends with lower accuracy in each seed, by 1.12 percentage points with shortcuts and 2.73 without. Alignment scores for credit assignment should be validated by intervention, not by correlation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.