PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 27, 2026Symmetry0 citationsOpen Access

Reinforcement Learning-Based Trajectory Planning for Dual-UGV Cooperative Carrying with Adaptive Target Entropy Regulation

View Full Paper
YZYue ZhouLGLijin GuoJLJ C Liu

Key Points

  • This research aims to improve trajectory planning for dual unmanned ground vehicles in complex environments, particularly for cooperative carrying.
  • Developed a reinforcement learning-based method with adaptive target entropy regulation.
  • Designed a local observation representation and cooperative reward function for distance maintenance and obstacle avoidance.
  • Implemented within a Multi-Agent Soft Actor-Critic framework under a centralized training and decentralized execution paradigm.
  • Generated safe coordinated trajectories with high relative-distance maintenance accuracy.
  • Simulation results showed significant improvements in trajectory planning effectiveness.
  • Demonstrated the adaptive adjustment of exploration intensity based on distance error.

Abstract

Trajectory planning for dual unmanned ground vehicles (UGVs) in cooperative carrying remains challenging in complex environments. The relative distance constraint imposed by the shared payload significantly increases the difficulty of cooperative trajectory planning. To address this issue, this paper proposes a reinforcement learning-based dual-UGV cooperative trajectory planning method with adaptive target entropy regulation. Specifically, a task-oriented local observation representation and a cooperative reward function are jointly designed for target guidance, distance maintenance, and obstacle avoidance, so that the learned policy can better satisfy the cooperative control objectives. Moreover, a target entropy regulation mechanism driven by relative distance error is incorporated into the maximum-entropy policy optimization process, enabling adaptive adjustment of the exploration intensity according to the current distance error. These designs are unified within a Multi-Agent Soft Actor-Critic (MASAC) framework under the centralized training and decentralized execution (CTDE) paradigm, forming a complete learning-based solution for dual-UGV cooperative carrying. Simulation results, including validation, comparative, and ablation studies, demonstrate that the proposed method can generate safe coordinated trajectories and achieve high relative-distance maintenance accuracy.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhou et al. (2026) studied this question.

synapsesocial.com/papers/6a16898b0c924ddd1bd583d2https://doi.org/10.3390/sym18050833
Ask AI
Helpful
Bookmark
Share
View Full Paper