acceptodds
Under review as a conference paper at ICLR 2027

Training Recipes Leak Through Behavioral Fingerprints in Text-to-Image Diffusion Models

Abstract

A deployed text-to-image diffusion model exposes generated images while keeping its fine-tuning procedure hidden. Training choices such as LoRA rank and scaling, learning rate, batch size, and training steps therefore appear inaccessible to a black-box user. We show that generated images can nevertheless reveal these choices. We present the first systematic study of whether fine-tuning hyperparameters of text-to-image diffusion models can be inferred from generated images alone. Given only black-box query access, we probe a target model with a fixed set of prompts and summarize its generations into a behavioral fingerprint using semantic, perceptual, and low-level image statistics. We then train predictors on fingerprints from shadow models with known recipes to infer the target model's hidden hyperparameters. Across 297 LoRA models, we recover five of six varied hyperparameters at rates significantly above chance. LoRA rank leaves the strongest fingerprint: we recover it with 85.4% exact accuracy and 100% accuracy within one grid step. When we consider all six predictions jointly, we recover the complete recipe exactly in 9.8% of cases, compared with a 0.21% uniform-chance baseline, a 48 increase over chance. We continue to recover rank significantly above chance across changes in model architecture, base-model weights, and fine-tuning data, and we find that controlling for adapter drift does not remove the signal. We also recover rank using only 96 generated images and continue to detect its fingerprint under substantial output perturbations. In contrast, LoRA dropout exhibits weak and classifier-dependent evidence of recovery. These results show that fine-tuning leaves persistent but uneven behavioral fingerprints in generated images, allowing black-box outputs to reveal parts of an otherwise hidden training recipe.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.