Visual perception in Unmanned Surface Vehicles (USVs) suffers from drastic lighting changes and missing texture features. These factors lead to depth scale drift and motion estimation bias. Moreover, existing multi-modal fusion models are computationally complex and unfit for resource-limited edge devices. To address these problems, a lightweight Radar–Vision Fusion (Geo-RVF) algorithm is proposed. To supplement spatial information, point clouds are projected to build sparse depth maps. A probability consistency-based depth correction module is designed to suppress water noise. This helps extract accurate geometric anchors to guide visual depth propagation. Subsequently, a Recurrent Autoregressive Network (RAN) fuses radar and image features in the temporal dimension. This resolves dynamic positional deviations caused by texture degradation and distant small targets. After real-time optimization, Geo-RVF achieves 23 FPS on the Jetson Orin NX. On a collected dataset, the method attains a mean average precision (mAP) 50–90 of 44.2% and a mean intersection over union (mIoU) of 99%, outperforming HybridNets and Achelous.
Zhou et al. (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: