MoGE-Bench: A Benchmark for On-device Diffusion Models
Abstract
Developing on-device image models requires considering practical user needs, task types and device resource constraints. However, existing benchmarks provide limited guidance on identifying capability gaps between on-device models and server-side models. We therefore introduce MoGE-Bench, a benchmark grounded in mobile user needs, comprising 64 subtasks and 1150 samples. It defines 16 operations, such as removal and replacement, and eight visual target categories covering content such as portraits and text. These form two orthogonal axes, with applicable combinations yielding explicit task coverage. To better guide model improvement, we propose MoGE-Eval, which introduces task-specific criteria covering instruction following, visual quality and detail preservation. It yields overall quality scores and capability profiles along the operation, target and criteria axes. We evaluate and deploy on-device models, and compare their output quality with server-side models. Our results show that this decoupled Operation × Target design provides finer-grained guidance for improving compact models. For example, it highlights the need to prioritize specific capability gaps (e.g., text-replacement) over broad application-level goals (e.g., improving text editing). We further construct a Quality-Deployability capability map using the deployment results. The map shows that current on-device models can achieve efficient generation or editing, but this deployment efficiency is not consistently accompanied by broad editing capability. Further progress requires improving editing quality and expanding task coverage within resource budgets, or exploring task-aware on-device–cloud routing. We will release the benchmark annotations, inference code and evaluation toolkit.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.