Semiconductor equipment production lines face challenges such as low equipment utilization and inefficient dynamic task allocation under the conditions of high product variety, small batch sizes, and highly variable demand. To address the difficulty in simultaneously achieving real-time responsiveness and optimality in resource scheduling, this paper models the production line as a Markov Decision Process (MDP) and proposes a dynamic resource scheduling algorithm based on Deep Reinforcement Learning (DRL). Specifically, the Deep Q-Network (DQN) serves as the core method. The state space includes equipment status, task queue, processing priority, and energy consumption information. The action space consists of allocation decisions for various tasks and equipment. The reward function integrates the minimization of task delay, equipment switching cost, and energy consumption. Experience replay and target networks are employed to enhance training stability. Simulation tests on a typical large-scale production line show that the DRL method achieves an average equipment utilization of 87.7%, an average order cycle time of 43.8 hours, and an optimal energy consumption value of 1285 kWh. The experimental results validate the efficiency and optimality of the proposed method in dynamically changing environments.
Yunpeng Fang (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: