acceptodds
Under review as a conference paper at ICLR 2027

Listening by the Right Amount: Calibrated Responses to User Requests in Text-Profile Recommenders

Abstract

Natural-language user profiles make recommenders scrutable: users can read what the system believes about them and correct it in words. A correction has a direction and an amount, yet evaluations and training objectives address only the direction, and nothing specifies how much a recommendation list should change. We measure thousands of requests, adding, withdrawing or negating a preference, on users the model never saw, against a principled yardstick for the amount: a user who states a preference should be treated like the reference users who already hold it. Fine-tuned rankers ignore requests (8% of a negated genre’s share removed on MovieLens-1M, 0% on tag-level attributes held out from training). Training them to obey fixes the direction but not the amount: after “I also love ⟨genre⟩ movies”, 98% of the top 20 carry that genre where reference users get 31%, and a withdrawal is executed as a negation; we show that an objective rewarding only the direction never prefers the right amount to a larger one, an error that success metrics cannot see. We propose calibrated edit distillation (CED). During training, a teacher given each request’s kind and attribute shifts the logits of the items it concerns by the log of the population odds ratio, which under a softmax choice model moves a neutral user’s odds exactly to those of the reference users; the model learns to produce this shift from the request’s text alone, and at the optimum only the requested items move. A 125M-parameter CED model removes negated genres completely and 85% of the share of held-out tags it never saw edited, and with a scale selected on validation users it lands within three points of the reference users on added and withdrawn genres and matches its teacher on genres. It follows 36 independently written phrasings once trained on 60 per kind and costs no measurable accuracy, without a separate parser or a generation step.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.