Monocular depth estimation is a fundamental task with broad applications in autonomous driving and augmented reality. While recent lightweight methods achieve impressive performance, they often neglect the interaction of mid-order semantic features, which are crucial for capturing object structures and spatial relationships that directly impact depth accuracy. To address this limitation, we propose MogaDepth, a lightweight yet expressive architecture. It introduces a novel Continuous Multi-Order Gated Aggregation (CMOGA) module that explicitly enhances mid-level feature representations through multi-order receptive fields. In addition, we present MambaSync, a global–local interaction unit that enables efficient feature communication across different contexts. Extensive experiments demonstrate that MogaDepth achieves highly competitive or superior performance on KITTI, improving key error metrics while maintaining comparable model size. On the Make3D benchmark, it consistently outperforms existing methods, showing strong robustness to domain shifts and challenging scenarios such as low-texture regions. Moreover, MogaDepth achieves an improved trade-off between accuracy and efficiency, running up to 13% faster on edge devices without compromising performance. These results establish MogaDepth as an effective and efficient solution for real-world monocular depth estimation.
Lin et al. (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: