• Propose of P-CoDETR for sewer pipeline automatic visual inspection from CCTV videos. • Use of RSwinTransformer to strengthen representation of complex geometries and spatial relationships. • Adaption of ATFL to dynamically emphasize hard samples, improving robustness and discriminative capability. • Achieving Highest mAP comparing to SOTA models among nine common pipeline defects. Employing computer vision-based models for intelligent pipeline defect detection offers a promising alternative to the conventional labor-intensive and time-consuming interpretation of CCTV videos. To better improve the capabilities of state-of-the-art object detection models addressing complex subterranean environment, this paper proposes a pipeline inspection model, P-CoDETR, with an enhanced Co-DETR architecture for defect detections under restricted field of view with imbalanced sample images. This model incorporates a Relation-Aware Self-Attention (RASA) mechanism within its backbone network, establishing an RSwinTransformer architecture to enhance modeling capabilities for complex geometries and spatial relationships. Concurrently, an Adaptive Threshold Focal Loss (ATFL) function is designed to dynamically adjust and focus gradients on challenging samples, thereby improving robustness and discrimination. Validation experiments have been conducted using a self-prepared large-scale dataset of 10,004 pipeline defect images of nine categories. The results demonstrate that compared to the original Co-DETR and mainstream YOLO/DETR series models, P-CoDETR achieves significant improvements in mAP and precision across defect types and model stability, exhibiting promising application prospects in replacing experienced engineers in pipeline automatic assessments. Notably, the average precision of total nine defect categories reaches 91.2% with remarkable improvements in detecting geometrically sensitive defects, such as deposition, corrosion and encrustation.
WU et al. (Sun,) studied this question.