acceptodds
Under review as a conference paper at ICLR 2027

What Individual Adapter Evaluations Miss: Composition-Only Failures in LoRA

Abstract

LoRA composition offers a way to reuse separately trained task adaptations, but evaluating the composed model raises a question that individual adapter checks leave open: does composition change which requests fail, even when average harm changes little? We study composition-only failures, where a composed response is labelled harmful on a request for which the base and every constituent alone are not. A request-level audit aligns base, member, and composition responses while keeping harmful compliance, refusal, and task performance distinct. Across several base models, we observe these failures in task-trained adapters and screened public adapter banks. The failures persist on additional requests, under two harm judges, and with regenerated responses. They are most pronounced when a composition includes an adapter that already produces some harmful responses alone: on average, harm stays close to that member's own rate while the set of failing requests shifts. They also occur in some compositions without such a member. Interventions expose a second gap. Retraining a member without safety data can lower the composition-only rate while increasing total harm, whereas adding refusal data lowers both and, in a tested summarization pair, retains the partner's task gain. Averaging composition weights removes most composition-only failures but leaves total harm above the base. Composed models should therefore be audited jointly for failure overlap, total harm, and retained adaptation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.