acceptodds
Under review as a conference paper at ICLR 2027

WorkSynth: Off-Policy Synthetic Data for Professional Artifact Agents

Abstract

Professional work benchmarks such as GDPval have begun to evaluate language agents not by isolated answers, but by their ability to produce complete, usable artifacts such as reports, spreadsheets, slide decks, analyses, and decision documents. This shift exposes a gap in current post-training practice. While recent reinforcement-learning methods have achieved strong gains on tasks with compact and verifiable outcomes, professional artifact generation is characterized by underspecified goals, heterogeneous formats, domain-specific conventions, and evaluation criteria that are only partially reducible to pass–fail rewards. We argue that part of this limitation reflects a distributional mismatch: general instruction-tuning data rarely teaches models the forms of situated work required to transform complex requests and supporting materials into polished deliverables. This paper studies whether such capabilities can be improved through off-policy post-training rather than expensive online agentic reinforcement learning. We construct synthetic professional work supervision that varies along three dimensions: task formulation, artifact structure, and revision-oriented feedback. Across GDPval-style evaluations, we examine whether models benefit more from additional final-output demonstrations, from critique-and-revision supervision, or from data targeting specific artifact conventions and professional genres. Our results show that synthetic work-oriented supervision improves instruction adherence, structural completeness, and deliverable quality, with the largest gains arising when synthetic data represents not only the desired final artifact but also the evaluative criteria that distinguish acceptable from incomplete work. These findings suggest that professional artifact agents may be limited less by a lack of generic reasoning ability than by insufficient exposure to the conventions, constraints, and revision practices of real-world knowledge work.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.