Harvesting SWE Benchmarks from Wild Vibe Coding Sessions
Abstract
Development teams need reproducible tasks and trustworthy evaluation signals that reflect their own engineering work to choose and improve coding agents. Meanwhile, as more developers start “vibe coding”, their everyday work leaves behind many in-the-wild coding sessions that record rich development context and could supply the raw material for such tasks. Yet these sessions are fragmented and loosely structured, making the valuable records hard to use directly. We present *SessionFarmer*, a framework that harvests these development records into executable software engineering (SWE) tasks. It uses developer–agent dialogue and code diffs to identify task boundaries, then brings together requirements that emerge across turns and sessions into coherent work episodes, each centered on one engineering goal. From each episode, it reconstructs a self-contained task specification. *SessionFarmer* also builds an executable environment and creates quality rubrics and unit tests for the task. A refinement loop is designed to reduce the chance that the tests accept incomplete implementations or reject valid alternatives. Using *SessionFarmer*, we construct CropChat with 242 tasks from public data and CropEnterprise from proprietary enterprise-level data. On CropEnterprise, evaluations by the original developers confirm the high fidelity and reliability of the reconstructed episodes and associated evaluation components. Furthermore, ablation experiments validate the effectiveness of the refinement stage in our construction pipeline. Together, these results show that *SessionFarmer* offers a promising path for unlocking everyday development records into reusable evaluation resources, with demonstrated applicability in real-world enterprise settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.