acceptodds
Under review as a conference paper at ICLR 2027

Output Coding in Multiclass Classification under Shared Bottlenecks

Abstract

In output coding for multiclass classification, each class is assigned a binary codeword, and the model learns the corresponding bits jointly from a shared feature representation. When the shared representation is constrained to be low-dimensional, a representation bottleneck, the binary prediction tasks compete for limited representational capacity, so the choice of code can affect which distinctions between classes are preserved. We study this effect under squared-loss training with a shared low-rank representation. Our goal is not to show that output coding universally improves classification, but to understand how the choice of code shapes the representation learned under this bottleneck. Even when two codes can represent exactly the same multiclass scoring rules, they can weight the training objective differently and therefore select different feature subspaces. We characterize this selected subspace and show that the resulting increase in classification error is proportional to the loss of Gaussian width, a geometric measure of the class structure removed by the projection. We also show that a linear decoder achieves the lowest classification error possible using only the retained representation. To show that conventional code properties do not determine which class information survives the bottleneck, for each bottleneck rank we construct pairs of binary codes with matched conventional code properties and binary-task difficulty, yet whose learned representations retain disjoint feature directions and have provably different multiclass risks. Moreover, the risk gap is at least of the improvement the Bayes classifier achieves over the best classifier that ignores the features, and this separation persists under small perturbations. Frozen-feature experiments further show that sensitivity to code choice can be substantial and task-dependent: coding improves over one-hot regression on Fashion-MNIST, but not consistently on CIFAR-100. These results show that conventional code properties alone do not determine which class information survives a shared representation bottleneck.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.