EmbTracker: Traceable Black-box Watermarking for Federated Language Models
Abstract
Federated language model (FedLM) training enables collaborative learning without sharing raw data. However, every client receives valuable model instances that may be redistributed without authorization, raising important intellectual property (IP) protection challenges. Most existing federated watermarking methods require white-box access or client-side cooperation, and provide only group-level ownership evidence rather than client-level attribution. Moreover, many methods are designed for image classification tasks and have limited utility for language models. We propose EmbTracker, a server-side framework for traceable black-box watermarking tailored for FedLMs. EmbTracker embeds a backdoor-based watermark that can be verified through API queries and assigns each client a distinct registered trigger through token-embedding replacement, avoiding per-client watermark retraining. Experiments on language and vision-language models across classification, question-answering and visual question-answering tasks show verification rates near 100% in most settings, primary-task changes typically within 1–2 percentage points, and resilience to fine-tuning, pruning, quantization, noise, and watermark overwriting under the evaluated conditions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.