FB-LoRA: Divergence-Free Gram Transport for Joint Bayesian Low-Rank Adaptation
Abstract
Low-rank adaptation enables parameter-efficient specialization of pretrained language models, yet reliable deployment requires calibrated uncertainty over the adapted weights. Directly applying Euclidean Langevin dynamics to low-rank factors injects isotropic noise, which disproportionately perturbs uninformative directions in the effective-weight space and degrades pretrained representations. To address this geometry mismatch, we introduce FB-LoRA, a Bayesian framework grounded in divergence-free Gram transport. We specify a proper generalized posterior over joint factor coordinates using a Gaussian prior and a bounded scoring loss. By preconditioning each factor’s drift and diffusion with a bounded spectral function of the opposite factor’s Gram matrix, the coordinate divergence of the mobility operator vanishes identically everywhere, eliminating state-dependent drift corrections without numerical approximations. This construction establishes a scale-invariant relative spectral allocation that concentrates stochastic exploration on dominant singular modes while guaranteeing localized finite-horizonconvergence. Across language understanding, mathematical reasoning, and safety benchmarks on LLaMA-3-8B, the resulting posterior ensemble maintains foundational capabilities and substantially suppresses nullspace jailbreak susceptibility. Furthermore, empirical predictive disagreement enables forward-only selective prediction and active information acquisition without test-time backpropagation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.