While vision-based crack segmentation automates construction inspection, leading to enhanced safety and long-term cost reduction, it is still challenged by complex crack geometries, cluttered environments, and computational constraints. This paper proposes VTECSeg, an encoder-decoder deep learning (DL) architecture combining a Vision Transformer backbone with a Convolutional Neural Network-based edge encoder to enable local crack detail preservation and global contextual reasoning. To further improve crack segmentation accuracy, a zero-shot Vision-Language Model-guided region proposal mechanism is introduced. Results demonstrate that VTECSeg outperformed several baseline architectures, achieving a precision of 87.32%, a recall of 84.21%, an F1-score of 85.74%, a mean Intersection-over-Union (mIoU) of 86.15%, and an Intersection-over-Union (IoU) of 79.5% on the DeepCrack dataset. This paper contributes to the body of knowledge by developing an edge-aware hybrid crack segmentation framework combined with a zero-shot vision-language-guided region proposal stage for high-fidelity and improved crack segmentation, thus advancing vision-based infrastructure inspection and automation. • Zero-shot VLM–SAM workflow for crack patch localization. • Combined text prompts and image encoding for spatial region proposals. • Fusion of vision transformers and edge-aware CNN features. • Pipeline outperforms existing deep learning models. • Supports real-time automated infrastructure inspection.
Poudel et al. (2026) studied this question.