acceptodds
Under review as a conference paper at ICLR 2027

MobileDev-Bench: A Benchmark for Issue Resolution in Mobile Application Development

Abstract

Large language models (LLMs) achieve strong performance on repository-level issue resolution, yet existing benchmarks focus on library-style repositories and rarely capture mobile applications, where fixes must build under framework-specific toolchains and often coordinate changes across source, resource, and build files. This paper introduces MobileDev-Bench, a benchmark of 407 verified issue-resolution tasks from 19 production Android Native (Java/Kotlin), React Native (TypeScript), and Flutter (Dart) applications, each paired with developer-written tests and a containerized build environment. Fixes modify 12.9 files and 334.6 lines on average, with 41% spanning multiple artifact types. MobileDev-Bench distinguishes tests that fail before the fix from those that become compilable only after the fix introduces required symbols. Under the single-pass Agentless pipeline, four LLMs (Claude Sonnet 4.5, GPT-5.2, Gemini 2.5 Flash, and Qwen3-Coder) resolve only 3.19%–4.18% of tasks with automated retrieval and at most 5.65% with oracle retrieval. With the interactive mini-SWE-agent, the best-performing model, GPT-5.2, resolves 12.53% of tasks, compared with 69.0% for the same model and agent on the SWE-bench Verified leaderboard. We release the benchmark tasks, evaluation harness, and containerized environments.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.