Fixed-wing unmanned aerial vehicle formation control confronts the dual challenges of achieving optimal performance amidst complex nonlinear dynamics and ensuring flight safety by constraining tracking errors. Existing reinforcement learning methods, though effective for optimal control, often overlook these critical safety constraints, which constitutes a serious shortcoming in safety-critical swarm operations where unmanned aerial vehicles confront highly nonlinear dynamics, unknown disturbances, and limited model knowledge, and these practical necessity drives us to synergistically integrate two previously separated techniques. To address this issue, this paper proposes a safe optimal control framework that synergistically integrates prescribed performance control with an actor–critic RL scheme. Specifically, a simplified actor–critic architecture is developed to derive a near-optimal controller for the coupled position–attitude dynamics without requiring an accurate model, thereby enhancing energy efficiency. Concurrently, the prescribed performance control is employed to transform the constrained formation error dynamics into an unconstrained system, guaranteeing that safety distances are strictly maintained. Lyapunov-based analysis proves that all signals in the closed-loop system are semi-globally uniformly ultimately bounded and that formation errors never violate the predefined performance boundaries.
Qiang et al. (Mon,) studied this question.