In this study, experimental deep reinforcement learning (DRL) control of a supersonic cavity flow is conducted for the first time at Mach 2, with the aim of mixing enhancement. A 4 5 pulsed-arc plasma actuator (PAPA) matrix with independently controlled columns and a supersonic hot-wire probe placed at the cavity midline serve as the flow disturber and state observer, respectively. The control law parametrised by a radial basis function network is executed on a field-programmable gate array at 5 kHz loop frequency. Results show that DRL is capable of finding a converged closed-loop control law in less than 10 s, and the resulting cavity velocity fluctuation is three times higher than periodic open-loop control. The control benefits earned by DRL increase with the number of activated columns, yet reduce with the cavity back-wall inclination angle. Using the same number of actuator columns, variable-formation actuation mode allows the DRL to find a more effective control with much less actuator power consumption, when compared with fixed-formation actuation mode. The final control law obtained by DRL can be interpreted as a threshold control conditional on the location of the state vector, and the improvement of total reward is ascribed to both the elevation of occurrence probabilities of high-reward clusters and the ubiquitous increase of the reward expectation at each cluster. Physically, mixing enhancement in the cavity flow is traced back to the thermal bulbs and shock waves produced by the PAPA, which induce a meandering motion of the shear layer.
Kong et al. (2026) studied this question.