Crack detection plays a pivotal role in ensuring the safety and stability of infrastructure. Despite advancements in deep learning-based image analysis, accurately capturing multiscale crack features in complex environments remains challenging. These challenges arise from several factors, including the presence of cracks with varying sizes, shapes, and orientations, as well as the influence of environmental conditions such as lighting variations, surface textures, and noise. This study introduces DAH-YOLO (Dynamic-Attention-Haar-YOLO), an innovative model that integrates dynamic convolution, an attention-enhanced dynamic detection head, and Haar wavelet down-sampling to address these challenges. First, dynamic convolution is integrated into the YOLOv8 framework to adaptively capture complex crack features while simultaneously reducing computational complexity. Second, an attention-enhanced dynamic detection head is introduced to refine the model’s ability to focus on crack regions, facilitating the detection of cracks with varying scales and morphologies. Third, a Haar wavelet down-sampling layer is employed to preserve fine-grained crack details, enhancing the recognition of subtle and intricate cracks. Experimental results on three public datasets demonstrate that DAH-YOLO outperforms baseline models and state-of-the-art crack detection methods in terms of precision, recall, and mean average precision, while maintaining low computational complexity. Our findings provide a robust, efficient solution for automated crack detection, which has been successfully applied in real-world engineering scenarios with favorable outcomes, advancing the development of intelligent structural health monitoring.
Fan et al. (2026) studied this question.