Multi-SWT-Bench: A Multilingual Benchmark for Reproduction Test Generation
Abstract
Reproduction test generation translates a natural-language issue description into executable tests that fail on the original code and pass after the issue is resolved, providing executable evidence for verifying candidate patches. However, existing benchmarks are constructed for individual programming languages, preventing a unified evaluation across diverse programming ecosystems. To address this limitation, we introduce Multi-SWT-Bench, a multilingual benchmark for reproduction test generation consisting of 1,963 instances across eight programming languages: Python, Java, TypeScript, JavaScript, Go, Rust, C, and C++. Based on this benchmark, we evaluate state-of-the-art LLMs using four representative methods (MSWE-agent, MOpenHands, Codex, and Claude Code) and analyze performance variations across language-specific subsets and factors associated with successful reproduction test generation. Our evaluation reveals a systematic language gap. Across every evaluated method and LLM, the Success Rate on Python exceeds the Success Rate across all languages, while the Success Rate on C++ is lower. This gap suggests that evaluations limited to Python can provide an overly optimistic view of reproduction test generation performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.