acceptodds
Under review as a conference paper at ICLR 2027

SCE-BENCH: FROM SYMPTOMS TO CAUSAL EXECUTIONS OF CONCURRENCY BUGS

Abstract

Concurrency bugs are hard to diagnose and understand: symptoms may not reveal their causes, and failures may depend on nondeterministic schedules. Yet existing LLM benchmarks either omit concurrency bugs or lack deterministic evaluation, limiting assessment of models' reasoning. We present SCE-Bench, an execution-grounded benchmark that tests whether models can turn symptom descriptions into machine-checkable causal executions. Each task tests this reasoning through four stages: diagnose whether a symptom is concurrency-induced, identify the mechanism by which concurrent interactions cause the failure, construct a workload, and submit an execution specification. The grader then runs each execution specification on buggy and repaired source revisions to check for the claimed failure and its absence after repair. SCE-Bench contains 66 positive tasks from disclosed kernel and userspace bugs across 12 mechanism families. Its 86 negative tasks comprise 66 repaired-source counterparts and 20 independent non-concurrency cases to test whether models distinguish actual concurrency bugs from false alarms. At pass@3, GPT-5.5 solved 9 of 66 positive tasks (13.6%), while only 10 were solved across all models. Particularly, they struggled most with the final stage of specifying an execution to trigger the bug. For the negative tasks, every model was less accurate on repaired-source controls than on independent non-concurrency cases.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.