Weed segmentation is a fundamental task in precision agriculture, essential for targeted intervention and sustainable farming. However, achieving accurate segmentation remains challenging due to the high visual similarity between weeds and crops, as well as the ambiguous, fine-grained boundaries often present in complex field environments. To address this, we present WS-DINO, a novel weed segmentation network built upon the DINOv2 vision foundation model. Our framework introduces two key innovations: (1) a Feature Prior Module that leverages a Canny-guided refinement process to extract and inject fine-grained cues related to weed texture, morphology, and boundaries into specific blocks of the Vision Transformer; and (2) a Spatial Feature Fusion Module that leverages convolutional layers to generate multi-scale spatial features, which are then fused with the semantically rich token features from DINOv2, effectively compensating for the Transformer’s limitations in capturing local spatial details. Comprehensive evaluation on the public PhenoBench dataset shows that WS-DINO achieves an mIoU of 88.67% and outperforms the evaluated benchmark methods. Moreover, on the challenging MotionBlurred dataset, WS-DINO reaches 88.75% mIoU, showing stable performance under motion blur and degraded visual conditions.
Zhou et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: