GraphForge: Training Working Agents with Graph-Anchored Workspace Synthesis
Abstract
Working agents need to read diverse files, coordinate tools, and produce deliverables. Training such agents requires tasks built on many real files with verifiable results, but few pipelines exist to synthesize this kind of data. Existing pipelines either generate files with models, which lack realism and diversity, or build tasks on real files without task-specific verifiers, leaving result quality unchecked. We introduce GraphForge, an evidence-graph based framework for synthesizing working agent training data. For each task seed, GraphForge first crawls the real files for its workspace and builds an evidence graph over the relations among these files. The graph is then compiled into task statements and rubrics with evidence anchors. An initial model rollout tests whether the task is executable, and a revise agent repairs any issues in the task statement and rubrics against the original files before the final trajectory is collected. The evidence anchors tell the judge exactly which files to consult when scoring each rubric. We fine-tune Qwen3.6-27B on 2,169 trajectories from GraphForge, bringing GDPVal to 1445.7 (+65.7), Workspace-Bench-Lite to 63.7 (+7.7), and SpreadsheetBench II to 24.0 (+13.7). Starting from the SFT model, we then apply offline rejection fine-tuning, sampling candidate trajectories, scoring them with the evidence-anchored rubrics, and keeping only the highest-scoring ones for another round of training. This brings further improvements over SFT on all three benchmarks. As the gain comes from training on the model's own rollouts, it suggests that the evidence-anchored rubrics provide useful selection signal.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.