Quantum reinforcement learning (QRL) is often evaluated under idealized, noiseless assumptions, yet realistic quantum devices inevitably introduce noise that can severely degrade performance. This paper improves the robustness of quantum deep Q-learning (QDQN) by redesigning the variational quantum circuit (VQC) used in its value-function approximator. Motivated by recent advances in quantum convolutional neural networks (QCNNs), we construct four QCNN-inspired VQC variants (Models A–D) by combining representative QCNN two-qubit building blocks with an explicit fully connected (all-to-all) layer. Using a 10-fold evaluation protocol at a fixed noise level p = 0.005, Model D achieves the best robustness, reducing the mean number of episodes required to reach a target reward from 1981 (baseline) to 1243. Under a stricter success criterion, Model D also doubles the empirically observed noise-tolerance boundary from 0.002 to 0.004. These results indicate that carefully chosen QCNN-style circuit components and connectivity can significantly improve the noise robustness of QDQN-like QRL agents.
Yu et al. (Tue,) studied this question.