acceptodds
Under review as a conference paper at ICLR 2027

Hybrid multi-objective reinforcement learning algorithm for hybrid flow shop with limited workers and time-of-use electricity tariffs

Abstract

With the increasing requirements of lower carbon emission, energy-related objectives have gained increasing attention in both industrial and management decision makers. This study proposed a Q-learning-driven hybrid reinforcement learning algorithm to solve the hybrid flow shop scheduling problem, wherein limited resources, worker energy consumption, and time-of-use (TOU) electricity tariffs are considered simultaneously. Three objectives are to be minimized, including makespan, total energy consumption, and TOU electricity pricing. First, a mathematical model based on the mixed integer linear programming (MILP) is formulated. A dynamic decoding approach is then developed to choose suitable machines and workers. Next, six variations of Nawaz–Enscore–Ham heuristics are proposed to enhance the initialization approach with diversity capabilities. Considering the multi-objective handling method, an adaptive-space-based crowding distance heuristic is embedded to select solutions resulted in higher global searching abilities. Additionally, two lemmas considering machine turn off and worker rest are designed to improve certain objectives. Furthermore, a Q-learning driven hyper-heuristic combining with VNS and IG is developed to balance the exploration and exploitation abilities. Finally, detailed experimental comparisons with state-of-the-art algorithms are tested to verify the efficiency and effectiveness of the proposed algorithm.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.