Rethinking the Harmonic Loss via Non-Euclidean Distance Layers
Abstract
Cross-entropy is the standard objective for training deep networks, yet it gives no inherent meaning to the learned weight vectors and drives them to grow without bound. The harmonic loss is a distance-based alternative that improves interpretability and mitigates grokking, but it has been studied only with the Euclidean distance, and never for its computational or energy cost. We replace that distance with a broad spectrum of metrics and evaluate the resulting distance-tailored harmonic losses on vision backbones and language models along three axes: model performance, interpretability, and sustainability. Across vision backbones, replacing the Euclidean distance is close to free: Bray-Curtis, cosine and Mahalanobis heads improve on Euclidean harmonic loss in most of the dataset-backbone settings we measure, and on the deeper backbones also on cross-entropy, at comparable measured emissions; where the harmonic family trails cross-entropy, on shallow backbones, the choice of distance still recovers most of the gap. On language modeling tasks, for smaller models, cosine-based harmonic losses improve gradient and learning stability, strengthen representation structure, and reduce emissions relative to cross-entropy and Euclidean harmonic loss. For a larger, modern decoder (SmolLM2), Euclidean harmonic loss attains the lowest perplexity and a Hamming-based harmonic loss the most structured representations at matched perplexity, at a higher training cost. On a pretrained 7B model, whose head spans only a few hundred classes, that cost disappears: harmonic heads match cross-entropy's accuracy while improving representation structure and out-of-distribution detection at no additional training cost. Our code is available at: https://anonymous.4open.science/r/rethinking-harmonic-loss/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.