Environmental perception in unstructured agricultural settings represents a fundamental challenge for autonomous orchard robotics. This study presents a real-time decision-level fusion framework integrating 32-channel LiDAR with a monocular camera for tree detection and localization in orchard environments. The perception system incorporates three key innovations: (1) triple-stage point cloud preprocessing utilizing region-of-interest segmentation, RANSAC-based ground removal, and intensity-based filtering to reduce computational load by 68%; (2) enhanced DBSCAN clustering with adaptive cluster merging and geometric shape filtering specifically optimized for tree morphology; and (3) decision-level fusion employing the Hungarian algorithm for optimal sensor association with semantic validation. A custom orchard dataset comprising 5000 annotated images was collected to train the YOLOv7 detection model, achieving 95.9% precision, 95.7% recall, and 75.2% mAP@0.5:0.95 on the test set. Experimental validation in a simulated orchard demonstrated superior performance: the fusion system achieved 97.80% precision and 96.50% recall, representing a 4.36% precision improvement over enhanced single-sensor approaches and a 29.44% gain vs baseline methods. Real-time processing capability was maintained at 41 ms per frame (24.4 FPS) with a 10-m effective perception range on moderate-cost hardware. Temporal robustness analysis across 100-frame sequences confirmed stable operation with 69% perfect detection, which is substantially higher than single-sensor methods. Field trials validated that the proposed LiDAR–camera fusion framework provides sufficient perception accuracy, robustness, and real-time performance to support practical autonomous navigation.
Lin et al. (2026) studied this question.