Cultivated land, as a vital resource for human sustenance, requires region-specific protection strategies worldwide. Semantic segmentation technology for agricultural land remote sensing imagery offers a scientific foundation and decision-making support for cultivated land protection through accurate identification and dynamic monitoring. In China, the fragmented distribution, small parcel sizes, complex terrain, and indistinct boundaries of cultivated land pose challenges to the intelligent interpretation of high-resolution remote sensing (HRRS) imagery. Conventional semantic segmentation methods often struggle to address these complexities. To address this issue, we propose a hybrid network called STAR-Net (Swin Transformer Auxiliary Residual Structure) for semantic segmentation of agricultural land in HRRS imagery whose encoder integrates a Global-Local Feature Fusion Module to effectively merge complementary information from both branches. A Multi-Scale Aggregation Module within the decoder facilitates the fusion of shallow spatial details and deep semantic cues, enhancing the model’s ability to discriminate objects at varying scales. Using the LoveDA dataset, we show that STAR-Net generates the highest Intersection over Union (IoU) on the “Barren” and “Forest”, achieving the improvement of 9.88% and 7.05% respectively, while delivering comparable IoU performance on other categories. Overall performance improved by 0.46% in mIoU compared to state-of-the-art models. Across all target categories, the method also achieves the greatest count of leading segmentation metrics.
Yang et al. (Fri,) studied this question.