SAFECOMPOSE: Predicting Unsafe LoRA Compositions Without Exhaustive Testing
Abstract
Adapter hubs let practitioners compose several LoRA modules on one open-weight base, yet each adapter is safety-screened only in isolation. We formalise composition screening as learning a set function over an adapter pool, separating a threshold crossing (safe parts, unsafe whole), its mechanism (distinct-adapter interaction versus summed merge dose), and genuine third-and-higher-order Mobius mass, which alone defeats low-degree predictors (Theorem thm:separation, a conditional result). We pair this with a screener that scores candidate compositions from adapter fingerprints and a calibrated accept/test/reject decision with a distribution-free false-negative guarantee (Theorem thm:ltt). Empirically, the picture is largely negative for interaction: across compositions of individually screened public adapters on a B base, none crossed the deployment threshold ( upper bound under the sampled law), and sub-threshold hazard is node-predictable with no resolved gain from interaction features; a deliberately constructed collusion is explained by summed merge dose, not a distinct-adapter mechanism. On a constructed benchmark whose unsafe compositions have every constituent individually safe at its coefficient, a coefficient-aware screener with a nonlinear dose term reaches AUROC (node-only ) and transfers to a held-out risky pair (–), while interaction architectures add nothing; an exploratory calibrated accept region holds out of sample at . The lesson is that merge configuration, not only the adapter file, is part of the safety object. Code and measured records: https://anonymous.4open.science/r/SafeCompose-F316
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.