Forward Similarity Does Not Guarantee Training Equivalence: Precision Placement in Differentiable Gram Solvers
Abstract
Changing the numerical precision of a differentiable Gram solver may leave predictions nearly unchanged while altering the computation used to learn its parameters. We study this distinction in controlled Spectral Koopman Attention (SKA) and an independent episodic ridge classifier, assigning FP32 or FP64 to separate solver stages without changing their exact-arithmetic definitions. Matched comparisons fix the model, input, and optimizer state; separately trained endpoints are evaluated through a common precision route. For SKA, we compare with the higher-precision solver route , a numerical reference rather than mathematical ground truth. Across 18 matched SKA validation comparisons from three training seeds, median relative discrepancies in logits, full-model gradients, and restored AdamW updates are , , and , respectively, each normalized by its corresponding reference norm. Prediction agreement can therefore miss differences in the learning computation. In a seven-seed matched classifier comparison, imposing externally prescribed ridge schedules shared across routes within each seed during training reduces the absolute between-route spread in post-training regularized-Gram conditioning by relative to learning ridge separately under each route. This shows that the observed conditioning separation depends on the training-time ridge policy. Precision choices in these solvers therefore warrant checks of parameter updates alongside predictions, with endpoint conditioning and task outcomes assessed separately.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.