PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 2, 20260 citationsOpen Access

Parameterized Reinforcement Learning with Route Guidance for Controlling Urban Road Traffic Networks

View Full Paper
EKEdwin M. KatakaTOThomas O. OlwalKDKarim Djouani

Key Points

  • The aim is to address limitations in traditional traffic control methods by proposing a new reinforcement learning approach for urban road networks.
  • Developed a Parameterized Deep Q-Network perimeter control scheme (P-DQNPC) for traffic management.
  • Trained and validated on synthetic macroscopic fundamental diagram data.
  • Evaluated performance using real-world traffic data from the San Francisco Bay Area Performance Measurement System.
  • Optimized both regional routing and continuous signal-timing actions.
  • P-DQNPC consistently outperformed traditional traffic control methods and state-of-the-art reinforcement learning models.
  • Achieved superior regulation of vehicle accumulation across regions.
  • Demonstrated enhanced performance in diverse and uncertain traffic conditions.

Abstract

Traditional macroscopic fundamental diagram (MFD)-based traffic perimeter metering control strategies rely on full knowledge of vehicle accumulation and inter-regional flow dynamics, assumptions that seldom hold in heterogeneous and highly variable real-world networks. Classical data-driven reinforcement learning methods face similar constraints, often converging slowly and exhibiting low sample efficiency when confronted with such complexities. Motivated by these limitations, this paper proposes a Parameterized Deep Q-Network perimeter control (P-DQNPC) scheme designed for multi-region urban road networks. The framework jointly optimizes discrete actions (regional routing choices) and continuous actions (signal-timing or flow-duration regulation) within a model-free learning structure. The approach is first trained and validated on synthetic MFD data to establish stable and interpretable policy behavior under controlled conditions. It is then transferred and further evaluated using real-world measurements from the Performance Measurement System—San Francisco Bay Area (PeMS-SF), a dataset collected from 18,954 loop detectors across the California State Highway System. PeMS-SF is selected due to its high spatial and temporal resolution, broad network coverage, and strong ability to capture realistic and diverse congestion patterns qualities that support both rigorous validation and generalization to other metropolitan regions. Experimental results show that P-DQNPC consistently outperforms state-of-the-art baselines, including deep deterministic policy gradient, deep Q-network, and No-Control schemes. The proposed method achieves superior regulation of regional accumulations and demonstrates enhanced robustness in large, heterogeneous, and uncertain urban traffic environments.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Kataka et al. (2026) studied this question.

synapsesocial.com/papers/69a52dbff1e85e5c73bf0d01https://doi.org/10.3390/futuretransp6020056
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Deep reinforcement learning-based adaptive traffic signal control in urban networks2026
  2. 2Coordinated Multi-Intersection Traffic Signal Control Using a Policy-Regulated Deep Q-Network2026 · 1 citations
  3. 3Combining multi-agent deep deterministic policy gradient and rerouting technique to improve traffic network performance under mixed traffic conditions2024 · 1 citations
  4. 4Deep Reinforcement Learning Technique for Traffic Metering in Connected Urban Street Networks2024
  5. 5Explainable reinforcement learning for improved traffic signal control2025 · 5 citations