acceptodds
Under review as a conference paper at ICLR 2027

GraphForge: Training Working Agents with Graph-Anchored Workspace Synthesis

Abstract

Working agents need to read diverse files, coordinate tools, and produce deliverables. Training such agents requires tasks built on many real files with verifiable results, but few pipelines exist to synthesize this kind of data. Existing pipelines either generate files with models, which lack realism and diversity, or build tasks on real files without task-specific verifiers, leaving result quality unchecked. We introduce GraphForge, an evidence-graph based framework for synthesizing working agent training data. For each task seed, GraphForge first crawls the real files for its workspace and builds an evidence graph over the relations among these files. The graph is then compiled into task statements and rubrics with evidence anchors. An initial model rollout tests whether the task is executable, and a revise agent repairs any issues in the task statement and rubrics against the original files before the final trajectory is collected. The evidence anchors tell the judge exactly which files to consult when scoring each rubric. We fine-tune Qwen3.6-27B on 2,169 trajectories from GraphForge, bringing GDPVal to 1445.7 (+65.7), Workspace-Bench-Lite to 63.7 (+7.7), and SpreadsheetBench II to 24.0 (+13.7). Starting from the SFT model, we then apply offline rejection fine-tuning, sampling candidate trajectories, scoring them with the evidence-anchored rubrics, and keeping only the highest-scoring ones for another round of training. This brings further improvements over SFT on all three benchmarks. As the gain comes from training on the model's own rollouts, it suggests that the evidence-anchored rubrics provide useful selection signal.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.