Are Small Language Models Ready for On-Device Agents? Benchmarking Small Language Models as Smart-Home Agents
Abstract
Small language models are increasingly used to build agents that run locally on resource-constrained devices. We introduce EdgeBench-Home, an executable benchmark for evaluating such agents through natural-language interaction with smart-home environments. Each episode specifies an environment state, an observation, user turns, typed tools, a resource budget, and an executable task target. The taxonomy covers seven task families spanning device control, multi-device selection, intent resolution, preference evolution, environment queries, automation, and robust tool interaction. A target-first construction pipeline generates candidate episodes and validates them through executable replay; the evaluator scores goal satisfaction, preservation, prohibited effects, and budget compliance while accepting semantically equivalent valid trajectories. The current core release contains 1,342 balanced bilingual episodes across five populated families. Our evaluation covers three model-size bands (≤1B, 1–3B, and 3–7B) and includes experiments on a device powered by the Rockchip RK3588 system-on-chip (SoC), with analyses of task capability and on-device resource behavior under different model and runtime settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.