This systematic review (PRISMA 2020) examines 89 studies—64 peer-reviewed articles and 25 arXiv preprints (2007–2026)—addressing the gap between AI research and operational predictive maintenance (PdM) deployment in complex manufacturing systems. Analyzing five thematic clusters in non-stationary and stochastic environments, we evaluated predictive performance and deployment readiness. Deep learning dominates remaining useful life (RUL) forecasting; however, 65.6% of studies employ weak or unclear validation protocols (Tier 0–1), lacking real-world robustness testing. Fault diagnosis increasingly integrates Edge-AI, yet Explainable AI (XAI) adoption remains scarce (15.6%), undermining industrial trustworthiness. No study reached operational field validation beyond temporal or cross-domain split, reflecting a systematic disconnection from deployed manufacturing systems. We introduce a novel Deployment Readiness Score (DRS) framework and identify critical barriers: data scarcity, environmental non-stationarity, computational constraints, and black-box model distrust. Recommendations include standardized temporal validation protocols, multi-site field studies, and architecture-integrated explainability. The 25 arXiv preprints (2024–2026) exhibit a mean DRS nearly three times that of the peer-reviewed corpus, signaling nascent convergence toward deployment-mature research. This review was not pre-registered.
Villa et al. (Fri,) studied this question.