Asking Back: Interaction-Layer Behavioral Transfer Through Distillation
Abstract
When one language model is distilled from another, more than answer content carries over: the student can also pick up how the teacher interacts with its users. We characterize this channel by planting two interaction-level markers in a teacher's responses, a trailing follow-up question and a declarative sentence of conditional advice woven into the answer, and then asking whether students trained on those responses reproduce them. We evaluate 63 student models spanning three pretrained families, seven data conditions, and three seeds, yielding 35,343 held-out responses. Both markers transfer even though the students are never given the marker instruction at test time. When every training response is marked, the interrogative marker appears in 41–81% of held-out student responses, versus 0–1.7% for students trained on the same prompts with unmarked responses. When only 19.34% of training responses carry the declarative marker, it appears in 4.34–8.32% of student responses. A five-model reference panel shows that unrelated instruction-tuned models already produce the interrogative marker in up to 16.93% of responses but produce the declarative marker in at most 1.60%, illustrating why ambient rate matters for interpreting a marker. In a 20-participant within-subject study, low-density marking left mean interaction-quality ratings similar across conditions. Response rewriting, surface filtering, and a green-list watermark run on the same data show which parts of the signal survive different transformations. We argue that interaction behavior is a measurable dimension of distillation, while keeping three questions apart throughout: whether a behavior transfers, whether it can be detected against a reference, and whether it can be attributed to a specific teacher.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.