acceptodds
Under review as a conference paper at ICLR 2027

Hackathon-24h: Evaluating Long-horizon Agents on Project-Scale Software Engineering

Abstract

Evaluating LLM-based agents on long-horizon software engineering tasks is highly challenging, as complex reasoning and execution behaviors unfold dynamically over time. To address this, we introduce Hackathon-24h, a novel evaluation protocol that shifts the paradigm from static measurement to continuous, trajectory-level process scoring. The protocol tasks agents with building project-scale systems software from scratch under a 24-hour wall-clock budget, requiring them to deliver runnable, incremental releases at regular intervals rather than withholding progress until the end. A decoupled architecture systematically snapshots and evaluates these updates against industrial conformance suites to capture continuous progress. Detailed analyses of these execution trajectories demonstrate the unique value of this continuous-delivery approach. By capturing intermediate dynamics, we uncover critical phenomena that final-state metrics completely miss such as hidden regression, delayed success, and temporal dynamics. We release our code at https://anonymous.4open.science/r/Hackathon-24h-anonymous/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.