ECO: Energy-Oriented Configuration Optimization for Attention–FFN Disaggregated LLM Serving
Abstract
Efficient large language model serving requires reducing energy while preserving latency and throughput. Attention–feed-forward network disaggregation exposes separate resource and operating choices, but their interactions create a large configuration space. Existing profile-based approaches require substantial offline measurements and need additional calibration data when deployed on new hardware. Untested configurations may also fail execution or violate service targets. We present Energy-Oriented Configuration Optimization (ECO), a framework for jointly tuning deployment and GPU operating settings within a limited measurement budget. ECO combines a calibrated stage-time and energy model with Gaussian-process corrections and cost-aware constrained Bayesian optimization. It separately models execution validity and service feasibility, selecting measurements according to expected energy improvement and GPU-time cost. The lowest-energy measured feasible configuration is then frozen for serving. Experiments with Qwen and DeepSeek cover A100 and A6000 GPUs. Across eight A6000 and four four-GPU A100 scenarios, ECO reduces serving energy by 44.4% and increases output token rate by 23.8% on average relative to Default while meeting service targets. Across eight A6000 scenarios, its selected feasible energy averages 33.1% below Bayesian optimization and 25.8% below genetic search.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.