This paper focuses on zero-sum game-based adversarial control for unmanned surface vehicles (USVs) against disturbances. Using the fully actuated system approach (FASA), the underactuated USV dynamics can be transformed into an equivalent fully actuated form. To solve the zero-sum game between the controller and disturbances, a reinforcement learning approach based on actor-critic architecture and policy iteration is proposed. The method avoids reliance on exact system dynamics by employing an off-policy learning scheme, where neural networks serve as function approximators for the cost function, control policies, and disturbance strategies. It is demonstrated that the iterative evaluation function gradually approaches the optimal value, while the combined weight matrix of all neural networks remains uniformly ultimately bounded. Simulations are conducted to demonstrate the efficacy of the proposed algorithm.
Xiang et al. (Tue,) studied this question.