Intelligent fault diagnosis of rotating machinery is essential for manufacturing reliability and predictive maintenance, yet deployment of deep learning models is limited by data scarcity: fault samples are rare, costly, and hazardous to obtain. Conventional synthetic data methods such as Generative Adversarial Networks and Variational Autoencoders often exhibit mode collapse, spectral distortion, and limited physical interpretability. This work presents MechaForge, a multi-strategy framework that employs Large Language Models (LLMs) as physics-guided generators for bearing fault time-series data. The approach is grounded in bearing kinematics, Motor Current Signature Analysis (MCSA), and the interpretation of in-context learning as implicit Bayesian inference. Within MechaForge, four progressively constrained tracks are defined: a real-data baseline, few-shot LLM mimicry, multi-stage semantic reasoning, and physics-guided generation with constraints on root mean square, kurtosis, and fault-band spectral energy. For direct benchmarking, conventional VAE- and GAN-based augmentation baselines are additionally evaluated under the same dataset split, synthetic-data budget, downstream CNN architecture, and evaluation metrics. Experiments on the Paderborn bearing dataset show that the Basic LLM track achieves the strongest performance under the present protocol (0.7862 accuracy, 0.7648 macro-F1), exceeding the added VAE and GAN baselines (both 0.7428 accuracy; 0.7202 and 0.7257 macro-F1, respectively), while a control experiment confirms that synthetic data provides discriminative structure rather than labeled noise. These results indicate the promise of LLM-based diagnostic augmentation under data scarcity in the present Paderborn setting, rather than a definitive demonstration of broad transferability across fault-diagnosis scenarios.
Zhang et al. (Wed,) studied this question.