PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 3, 2026Sensors0 citationsOpen Access

MCGC-Net: A Text-Enhanced Geometry-Consistent Network for UAV-Based Road Crack Detection

View Full Paper
ZOZhoujun OuSHShicong HeRBRongwei Bu

Key Points

  • This research aims to improve the accuracy of road crack detection using UAV images by utilizing multimodal data and advanced learning techniques.
  • Constructed a UAV-based multimodal road crack dataset with image-text annotations.
  • Introduced Multimodal Contrastive Semantic Gating (MCSG) to enhance visual feature learning.
  • Implemented Crack-Aware Slenderness Loss (CASL) to improve localization stability for slender cracks.
  • MCGC-Net achieved a significant increase in detection accuracy compared to traditional methods (exact metrics not specified).
  • Improved structural representation capabilities under complex road environments were observed, enhancing the differentiation between crack and non-crack areas.
  • Demonstrated robustness against challenges like blurred boundaries and irregular shapes in crack detection.

Abstract

With the rapid development of unmanned aerial vehicle (UAV) remote sensing and deep learning, road crack detection has become an important component of road condition assessment and intelligent road maintenance. However, accurately detecting cracks from UAV images remains challenging due to complex background environments, slender crack structures, blurred boundaries, and irregular crack shapes and orientations. Traditional methods that rely solely on visual information often struggle to achieve stable and accurate detection performance under these conditions. To address these challenges, this paper proposes a Multimodal Crack Geometry-Consistent Network (MCGC-Net) for high-precision road crack detection in complex road scenes. First, a UAV-based multimodal road crack dataset with image-text annotations is constructed. Specifically, crack-related textual descriptions are automatically generated from crack annotations using predefined semantic templates, which summarize crack morphology, spatial distribution characteristics, and structural properties. These semantic descriptions provide high-level semantic prior information for crack representation learning. Second, a Multimodal Contrastive Semantic Gating module (MCSG) is introduced to leverage automatically generated crack semantic descriptions and in-batch image-text semantic differences to guide visual feature learning, thereby improving the discrimination between crack and non-crack regions under complex background conditions. Furthermore, a Crack-Aware Slenderness Loss (CASL) is proposed to explicitly constrain slenderness consistency between predicted boxes and ground-truth boxes, improving localization stability for slender crack targets. In addition, a KAN-based Nonlinear Channel Attention mechanism (KAN-CA) is introduced to enhance feature representation capability for complex crack structures. Experimental results demonstrate that the proposed MCGC-Net effectively improves crack detection accuracy and structural representation capability under complex road environments. The proposed method provides a practical and reliable solution for UAV-based intelligent road crack detection.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ou et al. (2026) studied this question.

synapsesocial.com/papers/6a1fc530dee9eb8c0dce6964https://doi.org/10.3390/s26113487
Ask AI
Helpful
Bookmark
Share
View Full Paper