PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 18, 2026Transportation Research Record Journal of the Transportation Research Board0 citations

Multi-Intersection Traffic Signal Control With Deep Q Network Softmax Cross-Entropy Algorithm Based on Attention Mechanism

View Full Paper
XZX ZhangZXZiyao XiaLHLongji Huang

Key Points

  • The aim is to enhance traffic signal control efficiency using an improved deep reinforcement learning algorithm.
  • Developed a DQN softmax cross-entropy (DQN-SCE) algorithm for traffic signal control.
  • Utilized current phase and queue length as state representations.
  • Applied a multi-head self-attention mechanism to fuse state features.
  • Optimized the reward function by focusing on queue length.
  • Incorporated cross-entropy loss in both the target and action networks.
  • The DQN-SCE algorithm showed better performance in reducing average travel time than traditional methods.
  • Consistent improvements were seen compared to other reinforcement learning approaches.
  • The method exhibited enhanced convergence performance over existing DQN algorithms.

Abstract

With the continuous increase of urban traffic flow, the intelligence of traffic signal control (TSC) has become an important means to improve traffic efficiency. Among them, the deep reinforcement learning (DRL) algorithm Deep Q-Network (DQN) has been successfully applied to the field of TSC. We focus on the problems of complex state representation of existing traffic models, insufficient performance of DQN algorithm when using multilayer perceptron (MLP) as an action network, and over-estimation of Q-value leading to degradation of convergence performance. To mine the potential traffic state information from limited features and to improve the efficiency of the model, we propose a DQN softmax cross-entropy (DQN-SCE) TSC algorithm. First, the model uses the current phase and queue length as the state representation and optimizes the reward function only by the queue length. Second, a multi-head self-attention mechanism is used to fuse the state features. Finally, an improved DRL algorithm DQN-SCE is proposed; that is, we add cross-entropy loss of current actions for the target network and the action network to DQN. The experimental results based on CityFlow show that the TSC algorithm has better performance in the metric of average travel time compared with some traditional methods and reinforcement learning methods. The proposed algorithm still performs well compared with the traditional DQN algorithms and several improved algorithms for DQN.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2026) studied this question.

synapsesocial.com/papers/69e320fd40886becb654023fhttps://doi.org/10.1177/03611981261430695
Ask AI
Helpful
Bookmark
Share
View Full Paper