Automated detection of surface cracks in concrete structures is a fundamental requirement for effective structural health monitoring and preventive maintenance. Conventional computer vision and convolutional neural network‐based approaches often suffer from limited generalization under variations in illumination, surface texture, and crack morphology. To address these limitations, this study proposes a transfer learning framework based on Contrastive Language–Image Pretraining for concrete crack detection. A pretrained CLIP vision transformer with an input resolution of 224 by 224 pixels is employed as a frozen feature extractor, while a lightweight multilayer perceptron classifier is trained using the Adam optimizer with a learning rate of one times 10 to the power of minus three and a batch size of 32. The model is evaluated on the SDNET2018 dataset using a stratified train validation test split. Experimental results demonstrate a training accuracy of 89 point six percent, a validation accuracy of 94 point one percent, and a test accuracy of 88 point two percent, with a perfect recall and an F1 score of zero point 93. While the model achieves perfect recall and a high F1 score, the specificity remains comparatively lower, indicating a sensitivity‐biased behavior suitable for safety‐critical monitoring. These findings indicate that pretrained multimodal visual representations can effectively capture crack‐related texture patterns and offer a computationally efficient and reliable solution for real‐time concrete crack detection in structural health monitoring applications.
Md. Siam Ansary (Thu,) studied this question.