PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 7, 2026Journal of Computing and Information Science in Engineering0 citationsOpen Access

Deep reinforcement learning for helicopter assembly workshop scheduling considering the workers' operational proficiency

View Full Paper
ZZZihao ZhouYGYu GuoSHShaohua Huang

Key Points

  • To develop a scheduling method for helicopter assembly that considers workers' operational proficiency.
  • Proposed a deep reinforcement learning approach using Proximal Policy Optimization with Attention Mechanism.
  • Developed an assembly time prediction model considering worker skill levels and task criticality.
  • Formulated the scheduling problem as a Markov Decision Process (MDP) to manage scheduling dynamics.
  • The PPO-AM method outperformed traditional scheduling methods in complex assembly tasks.
  • Dynamic adjustments in the state and action spaces enhanced production efficiency.
  • The approach improved worker allocation flexibility in practical assembly scenarios.

Abstract

Abstract As a critical stage in helicopter manufacturing, the assembly process relies on effective scheduling to ensure both efficiency and quality. Traditional, experience-based scheduling methods are often inadequate for complex shop-floor environments, especially given the heterogeneity of worker skills and complex technological constraints. Therefore, this paper proposes a deep reinforcement learning approach based on Proximal Policy Optimization with Attention Mechanism (PPO-AM) to solve the helicopter assembly workshop scheduling problem (HASP) with consideration for worker skill proficiency. Firstly, Assembly time prediction model was established that integrates workers skill levels, task criticality, and the dynamic evolution of proficiency. This model provides a precise foundation for task-time estimation in assembly workshop scheduling. Based on this foundation, a shop scheduling model incorporating multiple process constraints was constructed. Furthermore, the state space, action space, and a reward function dynamically adjusted according to production progress were designed, thereby formulating the problem as a Markov Decision Process (MDP). Within this framework, a PPO-AM method incorporating a self-attention mechanism was proposed, and an assembly shop scheduling agent was developed based on this method to enable flexible and efficient workers allocation in practical scenarios. The PPO-AM method leverages the self-attention mechanism to assess the importance of state features, enabling the agent to adaptively focus on critical information, thereby enhancing its state awareness and policy generalization capability. the results demonstrated that PPO-AM exhibits superior performance and practical value in complex assembly scheduling tasks.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhou et al. (2026) studied this question.

synapsesocial.com/papers/69fbe357164b5133a91a2a3ahttps://doi.org/10.1115/1.4071848
Ask AI
Helpful
Bookmark
Share
View Full Paper