PolyAgent: A Multilingual Benchmark for API-Driven Mobile Device Control
Abstract
Small Language Models (SLMs) increasingly enable mobile agents that translate natural-language intent into executable device and application actions. While recent benchmarks have advanced multilingual function calling, stateful tool use, and interactive mobile-agent evaluation, these capabilities remain only partially integrated. In particular, localized multilingual intent grounding for executable mobile APIs under dynamic environment feedback is insufficiently characterized. We introduce PolyAgent, a multilingual diagnostic benchmark for API-driven on-device mobile control. PolyAgent contains 5,400+ expert-validated queries across 6 typologically diverse languages and 47 real-world mobile APIs. A culturally grounded synthesis and native-expert validation pipeline targets localized intent and argument grounding rather than direct translation alone. Integrated with ToolSandbox, PolyAgent evaluates interactive execution through Trajectory Precision and Goal Success Rate (GSR), combining action-level fidelity with environment-state verification. Experiments across lightweight open-source models reveal substantial model- and language-dependent variation under dynamic execution: Qwen models exhibit pronounced degradation for several non-English languages, whereas Gemma models maintain comparatively stable cross-lingual performance while achieving strong overall accuracy. These results highlight the limitations of static or English-centric evaluation for multilingual mobile tool use. PolyAgent provides a unified diagnostic framework for evaluating localized, state-aware, and resource-constrained mobile agents.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.