acceptodds
Under review as a conference paper at ICLR 2027

The Sins of the Designer: Weight Generators Transmit Judge Bias to Models That Never Saw a Biased Example

Abstract

Weight generators, such as hypernetworks and text-to-adapter systems, now produce LoRA adapters for downstream models that never see the generator's training data. If such a generator has learned a judge bias, does the bias reach the models it generates? We show that it does. We install a self-preference bias in a judge, distill it into a generator, and find that generated models inherit essentially all of it, across three scales of Qwen2.5, three of Llama-3, three bias attributes, and a publicly released text-to-LoRA generator that we update on a single task and then query on tasks it was never updated on. The transmission is directional rather than generic weight drift: only the bias direction transmits, a negated bias vector reverses the effect, and we measure the coherence between generation residuals and the bias gradient that a first-order bound says must be small, and find that it is. The transmitted bias distorts real rankings and survives clean finetuning, but it lives on a low-rank direction and can be projected out. Along the way we introduce g-PB, a generalized preference-bias metric that reduces to existing self-preference metrics as a special case.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.