The issue of collaborative task decision-making for autonomous underwater vehicles (AUVs) in partial observability is studied in this work. In view of the shortcomings of insufficient autonomous decision-making capabilities during multi-AUV collaborative tasks, a multi-agent deep reinforcement learning (MADRL) framework based on the multi-agent deep deterministic policy gradient (MADDPG) algorithm is proposed. The implementation of the centralised training with distributed execution (CTDE) framework enables each AUV to take global information as input during the training phase and to output actions based only on its own policy network during the execution phase. This allows each AUV to make decisions independently based on its local observations after training, minimizing the need for continuous communication and thus improving the system’s decentralized autonomous decision-making capability. In addition, a partially observable Markov decision process (POMDP) is designed for unknown marine environments to enable partially observable path planning for multi-AUVs. Simulation results demonstrate that the proposed approach can achieve autonomous obstacle avoidance and target assignment in multi-AUV scenarios. The proposed decision-making method enhances collaborative safety and facilitates the development of maritime mission operations.
Yu et al. (Wed,) studied this question.