EGSI ‐ PPO : An Evolutionary‐Guided Self‐Imitation Reinforcement Learning Framework for Autonomous Parking
Liang Hou et al.
What the paper says
With advances in autonomous driving technology, autonomous parking—an indispensable capability of intelligent vehicles—has emerged as a focal point for both academia and industry. To mitigate the slow convergence caused by sparse reward signals in parking tasks, this study introduces Evolutionary‐Guided Self‐Imitation Proximal Policy Optimisation (EGSI‐PPO), a novel algorithm that fuses the exploratory diversity of evolutionary strategies with the trajectory‐guided supervision of self‐imitation learning. The evolutionary component maintains policy diversity and enlarges the search space through population‐based parallel evolution, thereby enhancing global exploration, while the self‐imitation component transforms sparse rewards into dense supervisory signals using high‐return trajectories, simultaneously accelerating convergence and guiding the policy out of suboptimal traps. To balance parking accuracy, efficiency, and smoothness, a composite multi‐objective reward function is formulated, and a meta‐gradient weight‐balancing mechanism automatically adjusts the relative importance of each sub‐objective. In addition, action‐level smoothing and physical constraints are imposed at the policy output to ensure practical deployability. Experiments on the Webots simulation platform show that, compared with SAC, DDPG, and TD3, EGSI‐PPO delivers significant improvements in success rate, parking accuracy, and convergence speed. Ablation studies further confirm the individual contributions of the evolutionary component and the self‐imitation learning module. Overall, this work provides an efficient and robust deep reinforcement learning solution for autonomous parking and demonstrates the algorithm's potential in continuous control tasks characterised by sparse rewards.
Evidence weight
Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40
| F · citation impact | 0.50 × 0.4 = 0.20 |
| M · momentum | 0.50 × 0.15 = 0.07 |
| V · venue signal | 0.50 × 0.05 = 0.03 |
| R · text relevance † | 0.50 × 0.4 = 0.20 |
† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.