This paper presents on-policy and off-policy algorithms for the H∞ control of continuous-time mean-field stochastic Markov jump systems. Using online state and input data, these algorithms can learn control and disturbance strategies without the need for prior knowledge of system matrices. Under the standard assumptions that the system is mean-square stabilizable and detectable, we rigorously prove the monotonicity, boundedness, and convergence of the proposed iterative algorithms to obtain the unique stabilizing solution of cross-coupled generalized algebraic Riccati equations. Moreover, the off-policy algorithm features high data efficiency because the collected data can be utilized again after each iteration. Numerical simulations demonstrate the effectiveness of two algorithms and explicitly show that the off-policy algorithm achieves a faster convergence rate compared to its on-policy counterpart.
WANG et al. (Thu,) studied this question.