acceptodds
Under review as a conference paper at ICLR 2027

AskInfer-Bench: Profiles Are Evidence, Not Intent Labels, in Personalized Deep Research

Abstract

Personalized deep-research agents are often evaluated using criteria derived from the same persona supplied to the agent. This tests responsiveness, but it can silently turn a model's inference about a user into ground truth about the user's current task. We frame personalization under incomplete context as an information-acquisition problem: a profile is evidence about task intent, not the intent itself. AskInfer-Bench contains 15 tasks whose reference preferences are all human-provided, varies how much persistent profile evidence is initially visible, and independently controls clarification access. In the audited Prolific batch, ten adults supplied 107 raw statements, consolidated into 73 high-impact task-preference axes and 22 secondary preferences. A PDR-Bench-style GPT-5 generator produced 379 criteria for those ten tasks. Although its full lists attained 80.8% strict recall under an LLM-favorable semantic merge, recall fell to 39.7% after matching the 73-criterion human budget using the generator's own weights. The full-factorial system study contains 360 reports across two harnesses. Native clarification has near-zero average effects in the complete Gemini–DeerFlow matrix, whereas Open Deep Research with an Influence–Evidence–Ownership (IEO) question selector yields positive Ask–Off effects for three backbones under cold and RAW50 context but negative effects under RAW100. A separate five-pair dual-scoring diagnostic changes the average Ask effect from -0.2508 under human-grounded rubrics to +0.2924 under persona-derived rubrics and reverses the win/tie/loss conclusion for two pairs. This small diagnostic is not a population estimate or significance test; it illustrates why personalization benchmarks must separate using supplied information from deciding what still needs to be learned.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.