acceptodds
Under review as a conference paper at ICLR 2027

Signal Before Scale: User History and Inference Compute in LLM Personalization

Abstract

Personalizing a language model consumes two resources with unrelated cost structures: information about the user, and compute at inference time. Their marginal returns have been measured separately; their interaction has not, leaving open whether compute can buy back missing user signal. We cross a per-user signal axis against an inference-compute axis. The signal axis is the number k ∈ 0,...,32 of the user's own profile items retrieved into the prompt, manipulated within user rather than observed across users, removing the confound with activity level. The compute axis is selector-free: three model tiers and a chain-of-thought sampling budget over a frozen model. The design spans four LaMP tasks and 11,197 calls. On the three tasks where retrieved history moves the metric, eight items capture at least 96% of the total gain. At matched sample size, 39 of 71 signal contrasts have 95% bootstrap intervals excluding zero, against 10 of 82 for compute. On the citation-choice and rating-prediction tasks, the smallest model given eight of the user's own items outperforms the largest given none; on the latter it halves the error. The interaction is resolvable, and its sign depends on the task: model scale and user history substitute on rating prediction and complement on movie tagging, at four items and above. Any account of when inference compute pays for personalization must be task-conditional. A direct measurement explains why: at temperature 1.0, 38–67% of items return an identical answer on every sample, leaving self-consistency nothing to aggregate until history is in context. With no trained selector present, the shortfall lies in the candidate pool rather than in selection. Signal before scale: extra compute needs signal before it pays.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.