acceptodds
Under review as a conference paper at ICLR 2027

ECLIPSE: Orthogonal Subspace Projection for Triggerless Bias Injection in Diffusion Models

Abstract

The widespread adoption of modern multi-encoder text-to-image (T2I) diffusion models expands the attack surface for stealthy bias injection. This raises an important question in model security: In these complex, high-dimensional architectures, can attackers perform effective, triggerless bias injection while preserving structural fidelity? We introduce ECLIPSE, a framework for triggerless implicit bias injection. Unlike explicit backdoor attacks, implicit interventions avoid predefined triggers but can suffer from feature-level semantic entanglement, which causes conceptual bleeding. ECLIPSE addresses this challenge in two ways: (1) An encoder-specific Single-Stage Orthogonal Subspace Gradient Projection (OSGP) removes components of the accumulated LoRA gradients that align with dominant neutral feature directions before each optimizer update, concentrating optimization in their orthogonal complement. (2) LoRA-factor norm regularization limits excessive adapter changes while tuning a single target encoder, helping preserve non-target visual details without requiring a memory-intensive teacher model. Evaluations across the CLIP and T5-XXL text encoders of FLUX.1-schnell show that ECLIPSE achieves high manipulation success rates on fine-grained tasks, such as altering facial micro-expressions, as measured by a VLM judge calibrated against human annotations. It reduces interference between targeted semantics and non-target visual content while largely preserving the evaluated out-of-domain concepts, although limited changes remain. These findings highlight a potential security risk in the open-source generative supply chain and motivate architecture-aware defense mechanisms.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.