acceptodds
Under review as a conference paper at ICLR 2027

PersonalWorkEnv: Evaluating and Training Personalized AI Assistants with Synthetic Work Environments

Abstract

A personal work assistant can complete a task while violating a user-specific norm revealed in an earlier interaction. We call this failure mode the and introduce , a synthetic environment-generation pipeline for evaluating and training : remembering and applying previously revealed work norms to later tool-produced artifacts. Each environment combines a user persona, six simulated workdays, roughly twenty tasks, simple CLI tools, and ten progressively introduced profiles, including temporary norms that are later replaced. Deterministic validators and rubric-based soft checks distinguish surface task completion from profile compliance. We instantiate , a double-audited suite of fifty environments, and find a substantial personalization gap across seven proprietary and open models, with larger gaps on day 6 than day 1. As a baseline training use case, mixed-curriculum RL improves small Qwen3.5 models' all-check success. Conditional-compliance diagnostics show that the 9B gains primarily reflect better surface execution, whereas the 35B-A3B model also improves profile compliance among surface-successful tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.