acceptodds
Under review as a conference paper at ICLR 2027

OFFIDEX: Training Office Assistants to Clarify Before They Deliver

Abstract

The professionalism of a senior office worker rests on two abilities: producing deliverables that satisfy acceptance criteria, and eliciting, before any work starts, the constraints that decide whether a deliverable will be accepted. Large language model (LLM) assistants excel at the former on well-specified prompts but habitually answer first on the underspecified requests that dominate real workplaces, silently guessing deadlines, audiences, formats and data sources. We present OFFIDEX, an office assistant agent trained to clarify before delivering, and a Recursive Self-Purifying Experiential Reinforcement Learning framework that turns clarification into an optimizable objective without a single ground-truth answer. Three tiers of models cooperate: a frontier curator that purifies raw community knowledge into 125,982 task cards with hidden constraints and acceptance rubrics, mid-tier executors that act as a reveal-only-when-asked user simulator and a two-tier rubric/process judge, and a lightweight learner that improves through supervised fine-tuning and reward-filtered self-training. Every stage is guarded by a purification gate (data leakage, simulator fidelity, judge reliability, and a best-of-n gate that measures the learnable group-relative signal before any policy update). After purified fine-tuning the calibrated composite reward rises from −0.224 to +0.012 on validation prompts and stays positive (+0.020) on a fully held-out OfficeClarify, zero-clarification dialogues fall from 24% to 3–10%, and protocol violations fall from 2.06 per dialogue to zero. Five frontier models given the same clarify-then-deliver prompt, simulator and judge all score significantly below the trained assistant, answering first on 12–86% of requests: the habit is trained, not prompted. The loop’s diagnostics further show that hard hierarchical vetoes collapse open-ended office rewards onto a few fixed values and must be softened, and that reward is an inverted-U function of the number of clarifying questions, with one to two targeted questions optimal.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.