acceptodds
Under review as a conference paper at ICLR 2027

Same Labels, Different Formats: Auditing Conversational Personalization

Abstract

Giving a conversational protocol and an in-context learning (ICL) baseline the same labels does not control how those labels are presented. We audit this distinction in individual purchase prediction with frozen language models. On a 200-person cohort from Twin-2K-500, DeepSeek-V3.2 scores 0.477 F1 with block-format ICL (block ICL) and 0.699 with a conversational outcome-feedback protocol. Presenting the same seven labelled examples as a fixed dialogue within one request yields 0.697, without generated context turns or feedback on model predictions; the paired protocol-minus-dialogue difference is +0.002 (95% participant-bootstrap interval [-0.031, 0.033]). A disjoint 200-person cohort reproduces the pattern (0.429, 0.703 and 0.705 for block ICL, dialogue ICL and the protocol). The direction depends on the model: Qwen-2.5-72B scores 0.705, 0.668 and 0.657; a wording-matched single-message control shows that the lever is message format (dialogue minus control +0.311, [0.241, 0.382]). Re-running three conditions changes F1 by less than 0.01. Further controls show that the conversation’s advantage is carried by the participant’s own history rather than the demographic profile, and that accuracy and prediction propensity must accompany F1; a nine-model panel, two slate-choice panels and an acknowledgment intervention set the audit’s scope. Across the three models run under matched conditions, a three-call non-interactive baseline performs at or above the 49-call conversational protocol, which shows no resolved advantage on any of them; presentation, not only labels, is therefore the baseline against which claims about interaction should be judged.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.