High-precision identification and three-dimensional (3D) positioning of cabbage plants across their entire growth cycle are fundamental prerequisites for automated agricultural management. To overcome field challenges like extreme morphological variations, severe leaf occlusion, and bounding box jitter, we introduce a camera-LiDAR fusion perception system. First, an advanced SCEW-YOLOv8 architecture is proposed, sequentially integrating SPD-Conv downsampling, a C2f-CX global feature enhancement module, an EMA cross-space attention mechanism, and the WIoU v3 loss function. Evaluated on a comprehensive whole-growth-cycle cabbage dataset, the model achieves 95.8% mAP@0.5 and 90.8% recall with a real-time inference speed of 64.2 FPS. Furthermore, a visual semantic-driven camera-LiDAR fusion ranging algorithm is developed. Through rigorous spatiotemporal synchronization and cascaded outlier filtering, the integrated system achieves millimeter-level 3D localization within the typical 1.0–2.0 m operating range of agricultural robots. It maintains a Mean Absolute Error (MAE) of only 1.45 mm in the longitudinal direction at a stable processing throughput of 20 FPS. Compared to traditional pure vision depth estimation, this heterogeneous fusion approach achieves a remarkable 96.3% reduction in spatial positioning error at extended distances, fundamentally eliminating depth degradation caused by complex illumination. Ultimately, this system provides a highly robust, full-cycle geometric perception framework for the autonomous management of open-field green cabbage.
Han et al. (Fri,) studied this question.