Efficient Charging Strategies for Electric Trucks using Actor-Critic Reinforcement Learning with Tree Ensembles
Abstract
Long-haul electric trucking requires charging strategies that balance delivery efficiency against battery health, under uncertainties such as stochastic energy consumption. We formulate this as a profit-maximizing sequential decision problem: At each charging station, an agent chooses how long to charge and at what power, trading delivery revenue against electricity cost, driver wages, and battery degradation. The environment is inherently non-stationary since the agent's own decisions determine how battery capacity evolves over time. Addressing this problem well calls for methods that remain effective under drifting dynamics and can run on constrained, embedded, in-vehicle compute. To address this challenge, we propose tree ensembles combined with reinforcement learning. Tree ensembles excel on tabular low-dimensional data typical of such control problems, but remain largely unexplored for actor-critic reinforcement learning. The only existing method, which grows ensembles without bound during training, incurs continuously increasing memory and compute costs. As an alternative, in this paper we introduce Tree Ensemble Reinforcement Learning (TERL), which instead fixes the ensemble size. To update the model, TERL periodically rebuilds trees with off-policy reinforcement learning, and refits leaf values with on-policy methods between rebuilds. This requires no incremental tree-growing interface, while still letting the ensemble continuously learn and adapt to non-stationary dynamics. We instantiate TERL with XGBoost and demonstrate its competitiveness against existing methods. We evaluate both on standard benchmark problems and our electric-truck charging environment, offering a realistic testbed for EV charging optimization and a general, resource-efficient framework for tree ensemble reinforcement learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.