acceptodds
Under review as a conference paper at ICLR 2027

PACE: Rethinking On-Device LLM Efficiency From Token Throughput to Timely Correctness

Abstract

On-device LLM systems are often optimized for token throughput or energy per token, but these metrics do not capture whether a request is answered correctly and on time. We present PACE, a task-aware framework for evaluating on-device LLM efficiency through timely accuracy (i.e., the fraction of requests answered correctly before a deadline) and total request energy. PACE studies how deployment, execution mode, and generation budget jointly affect task completion. Across three LLMs and three deployment paths on two real mobile phones, we find that similar token-level efficiency can hide large differences in task-level cost: two deployments with only a 1.01% difference in energy per token differ by 2.14× in energy per timely correct answer at a 10-second deadline. We further find that faster execution matters only when it helps correct responses meet the deadline, while larger generation budgets matter only when the extra tokens improve answer correctness. We then introduce PACE-SELECT, a simple calibration-based rule that chooses the lowest-energy configuration satisfying a target timely accuracy. Our results show that the most efficient configuration depends on the task requirement and deadline, and that on-device LLM efficiency should be evaluated by timely task completion rather than token efficiency alone.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.