PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 14, 2026Expert Systems0 citations

EGSI ‐ PPO : An Evolutionary‐Guided Self‐Imitation Reinforcement Learning Framework for Autonomous Parking

View Full Paper
LHLiang HouFZF ZhangJCJie Cao

Key Points

  • This research aims to develop a new algorithm to improve the performance of autonomous parking systems.
  • Introduces EGSI-PPO algorithm combining evolutionary strategies and self-imitative learning.
  • Uses population-based evolution to enhance policy diversity and global exploration.
  • Implements a multi-objective reward function to balance accuracy, efficiency, and smoothness.
  • Carries out simulations on the Webots platform for performance evaluation against existing algorithms.
  • EGSI-PPO shows significant improvements in success rate compared to SAC, DDPG, and TD3.
  • Demonstrates enhanced parking accuracy and convergence speed during simulations.
  • Ablation studies reveal the beneficial individual contributions of both evolutionary and self-imitative components.

Abstract

ABSTRACT With advances in autonomous driving technology, autonomous parking—an indispensable capability of intelligent vehicles—has emerged as a focal point for both academia and industry. To mitigate the slow convergence caused by sparse reward signals in parking tasks, this study introduces Evolutionary‐Guided Self‐Imitation Proximal Policy Optimisation (EGSI‐PPO), a novel algorithm that fuses the exploratory diversity of evolutionary strategies with the trajectory‐guided supervision of self‐imitation learning. The evolutionary component maintains policy diversity and enlarges the search space through population‐based parallel evolution, thereby enhancing global exploration, while the self‐imitation component transforms sparse rewards into dense supervisory signals using high‐return trajectories, simultaneously accelerating convergence and guiding the policy out of suboptimal traps. To balance parking accuracy, efficiency, and smoothness, a composite multi‐objective reward function is formulated, and a meta‐gradient weight‐balancing mechanism automatically adjusts the relative importance of each sub‐objective. In addition, action‐level smoothing and physical constraints are imposed at the policy output to ensure practical deployability. Experiments on the Webots simulation platform show that, compared with SAC, DDPG, and TD3, EGSI‐PPO delivers significant improvements in success rate, parking accuracy, and convergence speed. Ablation studies further confirm the individual contributions of the evolutionary component and the self‐imitation learning module. Overall, this work provides an efficient and robust deep reinforcement learning solution for autonomous parking and demonstrates the algorithm's potential in continuous control tasks characterised by sparse rewards.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hou et al. (2026) studied this question.

synapsesocial.com/papers/69b4fc0eb39f7826a300c9d8https://doi.org/10.1111/exsy.70240
Ask AI
Helpful
Bookmark
Share
View Full Paper