Ship detection in Synthetic Aperture Radar (SAR) imagery remains challenging due to complex backgrounds, scale variations, and limited semantic discrimination in conventional detectors. To address these critical challenges, we propose SAR-SwinX (SAR Swin Transformer-enhanced YOLOX): a novel hybrid lightweight one-stage detection model composed of anchor-free Exceeding You Only Look Once (YOLOX) as a baseline and enhanced with Swin transformer modules based on cross-stage partial connections (CSP) to improve contextual representation. This hybrid design combines the local feature extraction strengths of Convolutional Neural Networks (CNNs) with the global semantic modeling of visual transformers, enabling effective multiscale ship detection in cluttered maritime scenes. Extensive experiments conducted on two public SAR datasets, including SSDD and HRSID, consistently validate the superiority of SAR-SwinX over the baseline YOLOX-s and existing state-of-the-art methods. A key result of our approach is that SAR-SwinX improves mAP@50:95 by 2.03% and 3.46%, enhances recall by 0.73% and 0.52%, and boosts the F1-score by 0.56% and 0.16% for SSDD and HRSID, respectively. These results highlight SAR-SwinX as an efficient and robust solution for SAR ship detection in complex environments, with favorable computational efficiency.
Yehia et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: