PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 9, 2026Automation in Construction0 citationsOpen Access

VTECSeg: Edge-aware hybrid CNN-vision transformer network with zero-shot vision-language-guided region proposals for crack segmentation

View Full Paper
OPOscar PoudelXHXi HuRARayan H. Assaad

Key Points

  • The aim is to enhance crack segmentation accuracy in construction inspection using a new hybrid model.
  • Developed an encoder-decoder architecture combining Vision Transformer and CNN features.
  • Introduced a zero-shot Vision-Language Model-guided region proposal mechanism.
  • Evaluated the model on the DeepCrack dataset.
  • Achieved a precision of 87.32%, recall of 84.21%, and F1-score of 85.74%.
  • Recorded a mean Intersection-over-Union (mIoU) of 86.15% and IoU of 79.5%.
  • Outperformed several baseline architectures in crack segmentation.

Abstract

While vision-based crack segmentation automates construction inspection, leading to enhanced safety and long-term cost reduction, it is still challenged by complex crack geometries, cluttered environments, and computational constraints. This paper proposes VTECSeg, an encoder-decoder deep learning (DL) architecture combining a Vision Transformer backbone with a Convolutional Neural Network-based edge encoder to enable local crack detail preservation and global contextual reasoning. To further improve crack segmentation accuracy, a zero-shot Vision-Language Model-guided region proposal mechanism is introduced. Results demonstrate that VTECSeg outperformed several baseline architectures, achieving a precision of 87.32%, a recall of 84.21%, an F1-score of 85.74%, a mean Intersection-over-Union (mIoU) of 86.15%, and an Intersection-over-Union (IoU) of 79.5% on the DeepCrack dataset. This paper contributes to the body of knowledge by developing an edge-aware hybrid crack segmentation framework combined with a zero-shot vision-language-guided region proposal stage for high-fidelity and improved crack segmentation, thus advancing vision-based infrastructure inspection and automation. • Zero-shot VLM–SAM workflow for crack patch localization. • Combined text prompts and image encoding for spatial region proposals. • Fusion of vision transformers and edge-aware CNN features. • Pipeline outperforms existing deep learning models. • Supports real-time automated infrastructure inspection.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Poudel et al. (2026) studied this question.

synapsesocial.com/papers/69fecf16b9154b0b82876394https://doi.org/10.1016/j.autcon.2026.106998
Ask AI
Helpful
Bookmark
Share
View Full Paper