LU-500: A Logo Benchmark For Concept Unlearning
Abstract
Concept unlearning is increasingly used to limit protected or unsafe visual concepts in text-to-image models, yet existing evaluations mostly study targets that dominate an image rather than localized company logos. Logos create a distinct failure mode: a small mark can carry the protected concept, remain recognizable from partial visual evidence, and be triggered implicitly by products, storefronts, packaging, or advertisements. We introduce LU-500, a Fortune Global 500 logo-unlearning benchmark containing 9,584 human-verified text-query and logo-image pairs across an explicit track (LUex-500) and an implicit contextual track (LUim-500). Our multi-grained protocol evaluates local removal and non-logo preservation, sup- plements detector outputs with fixed-region and full-output audits, and validates both components against blinded human judgments. Representative inference-time and compatible fine-tuning methods show an erasure–preservation tension under their tested global interventions. This pattern persists with zero-fill-free masked metrics and three seeds, whereas a detect-then-inpaint control attains 78.2% joint success, demonstrating that the benchmark can reward a localized solution rather than mechanically inducing a trade-off. A small SD3/FLUX study yields the same broad operating regimes, providing limited portability evidence. We also analyze ProLU as a prompt-space diagnostic: semantic rewriting can reduce logo triggers, but it is neither weight-level unlearning nor evidence of latent-space entanglement.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.