Signs Beat Floats: Low-Rank Double-Binary Adaptation for Storage-Efficient Deployment
Abstract
Task-specific LoRA adapters can be stored and exchanged independently of a shared frozen language model, but their floating-point factors remain a substantial deployment cost. We introduce LoRDBA, a low-rank double-binary adapter that represents both factors with binary signs and lightweight channel-wise scales. A post-training factorization initializes the signs, after which our default QAT Freeze procedure adapts only the scales; QAT Full additionally updates latent carriers. We provide a conditional finite-sample reconstruction bound for a sign-noise model. Across three-seed math-reasoning evaluations, LoRDBA achieves favorable compact accuracy-storage trade-offs on LLaMA-2-7B and 13B, with model-dependent margins on Mistral-7B. At approximately MiB logical payload on 7B, Freeze improves GSM8K/Minerva means by points over compact fp16 LoRA with equal total dataset passes. On 13B, it exceeds an SVD rank-2 control initialized from the same source and given the same two-epoch recovery by points at approximately MiB. Separately, packed execution reduces resident adapter memory by , and adapter-local recomputation brings Freeze training to the fp16-LoRA peak-memory level at higher step time. These findings support storage-efficient adapters produced on server GPUs, not inference speedup or demonstrated on-device training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.