SourceWorldBench: Can Coding Models Predict Test and Profiler Outcomes?
Abstract
Coding benchmarks increasingly focus on realistic repository-level software engineering tasks approached via agentic harnesses, where the models are allowed to explore repository contents and execute code. Supporting code execution for agentic harnesses requires sophisticated infrastructure, so code world modeling, i.e., simulating the program's execution, has recently gained traction. Performance optimization, in particular, is a common software engineering task that often requires high-load or long-running workloads and could substantially benefit from simulation. However, while simulating the logical code state during execution and end-to-end performance optimization are widely studied, it remains unknown whether models can reliably, say, simulate a profiler. Thus, we introduce SourceWorldBench, a repository-level benchmark comprising six tasks from various aspects of modeling software performance: test outcome prediction, execution time and traced memory estimation, and runtime and memory hotspot localization. We derive instances from an end-to-end performance optimization benchmark SWE-fficiency, augment them with agent-generated patches, and execute test suites and performance workloads to construct the benchmark labels. We evaluate six proprietary and open-weight models under four context collection baselines, ranging from no retrieval to privileged execution-informed context. For all six tasks, the best configurations outperform simple task-specific heuristics marking the possibility of progress, while many of the considered configurations do not, highlighting a substantial room for improvement. For example, for the task of test outcome prediction, Claude Opus 5 and GPT-5.6 Sol score 45.1 and 41.5 Macro F1 respectively, while a random coin flip scores 49.5. These findings establish SourceWorldBench as a challenging testbed for studying and improving optimization-adjacent code world modeling capabilities in isolation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.