The field of intelligent transportation on inland waterways is experiencing rapid growth, driven by the global pursuit of enhanced waterway safety, operational efficiency, and environmental sustainability. In real-world autonomous operation scenarios of unmanned surface vehicles (USVs), image-based 2D object detection methods are insufficient to meet the demands of 3D environmental modeling and accurate perception of dynamic objects. Existing 3D perception systems for USVs depend heavily on precise sensor calibration. However, projection offsets between point clouds and images—caused by water surface fluctuations and complex outdoor environments—hinder the practical deployment of these methods. To address these limitations, we propose a weak calibration multi-modal 3D object detection algorithm based on cross-view fusion, termed RCF-Free (Radar-Camera Fusion, Free from precise calibration). Inspired by autonomous driving solutions, we design a Triple-Path Cross-View Fusion module that achieves high-quality cross-view feature fusion without requiring accurate calibration parameters, while simultaneously detecting complete bird’s-eye view (BEV) bounding boxes. We further enhance the spatial layout comprehension of the visual branch through a Mobile Self-Attention Module (MAM) and effectively encode sparse point cloud features in BEV space using a dedicated BEV-Point feature encoder. Additionally, we reconstruct and introduce two water-related 3D object detection datasets, FloW-BEV and WaterScenes-BEV. Experimental results demonstrate that RCF-Free achieves mAPBEV50 scores of 60.5% and 69.3% on the FloW-BEV and WaterScenes-BEV datasets, respectively, showing the effectiveness in water surface object detection. Moreover, on the DAIR-V2X-I dataset for autonomous driving scenarios, the model attains mAP3D50 scores of 73.3%, 61.2%, and 61.2% across three task difficulty levels, illustrating strong cross-domain generalization capability.
Li et al. (Wed,) studied this question.