Pairs or Bits: Two Repairs for the Binarized Convolution
Abstract
A binarized convolutional network computes each unit as a weighted count of input bits, one-bit activations against ternary weights. It deploys as a multiplier-free circuit and pays in accuracy, almost all of it charged to the activations and not the weights. Two repairs act inside the unit. One widens the activation to two or three bits, which widens every accumulator and the memory between layers. The other keeps one bit and lets each unit multiply its inputs in pairs as well as summing them, for four integer multiplies. The literature has developed the first. We develop the second and price both on what a circuit carries, logic area and the activation bits each layer passes on. We prove that the pair unit still compiles exactly, that neither repair's class contains the other, and that a rank condition on the pair-coefficient matrix, minimized over its diagonal, decides which degree-2 polynomials a rank- unit realizes. Under one recipe for every model the pair term returns of the that binarization costs on CIFAR-10 at equal parameters, and at equal logic. On that dataset and four more it returns between and while a second activation bit returns between and , so which repair pays depends on the task. On two of the five the pair term leads while storing a third of the activation bits where the bit stores twice as many. It also reaches its plateau in a little over half the epochs, because a pair term gives the integer its threshold compares far more values to take. Neither the stored bits nor the speed depends on the learning rate. The accuracy gap does, and a tuned baseline closes it. Where activation bits are affordable, add them. Where memory or interconnect binds rather than arithmetic, add the pair term instead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.