This study addresses the challenge of detecting white grape bunches (Vitis vinifera L.) in high-density vineyard canopies, a critical task for precision viticulture and yield estimation. Traditional statistical and image-processing methods have struggled to cope with occlusion issues. In this work, more than 200 field RGB images were collected at La Bergonza (Toledo, Spain) and expanded through data augmentation. Several preprocessing strategies were evaluated to enhance bunch visibility. Different convolutional neural network (CNN) architectures were compared, with YOLOv8 outperforming Mask R-CNN in terms of both accuracy and efficiency. YOLOv8, trained for up to 100 epochs on equalized and augmented datasets, achieved outstanding performance, with 84.9% precision, 72.6% recall, and an mAP@0.5 of 83%, far surpassing Mask R-CNN (17% precision and 26% recall). The model successfully detected partially occluded grape bunches, including some that were not visible to human experts, and outperformed previous studies that relied on controlled backgrounds or artificial lighting. The results demonstrate that combining RGB equalization with data augmentation significantly improves detection performance. These findings highlight the potential of deep learning and low-cost RGB imaging systems to enable automated and scalable solutions for yield estimation and canopy analysis. In conclusion, YOLOv8 emerges as a promising tool for accurate grape bunch detection under real field conditions, effectively overcoming previous technological limitations.
Fuentes et al. (Wed,) studied this question.