PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
July 26, 20244 citationsOpen Access

UniM-OV3D: Uni-Modality Open-Vocabulary 3D Scene Understanding with Fine-Grained Feature Representation

View Full Paper
QHQingdong HeJPJinlong PengZJZhengkai Jiang

Key Points

Key points are not available for this paper at this time.

Abstract

3D open-vocabulary scene understanding aims to recognize arbitrary novel categories beyond the base label space. However, existing works not only fail to fully utilize all the available modal information in the 3D domain but also lack sufficient granularity in representing the features of each modality. In this paper, we propose a unified multimodal 3D open-vocabulary scene understanding network, namely UniM-OV3D, aligning point clouds with image, language and depth. To better integrate global and local features of the point clouds, we design a hierarchical point cloud feature extraction module that learns fine-grained feature representations. Further, to facilitate the learning of coarse-to-fine point-semantic representations from captions, we propose the utilization of hierarchical 3D caption pairs, capitalizing on geometric constraints across various viewpoints of 3D scenes. Extensive experimental results have demonstrated the effectiveness and superiority of our method in open-vocabulary semantic and instance segmentation, which achieves state-of-the-art performance on both indoor and outdoor benchmarks such as ScanNet, ScanNet200, S3IDS and nuScenes. Code is available at https://github.com/hithqd/UniM-OV3D.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

He et al. (2024) studied this question.

synapsesocial.com/papers/68e5ee87b6db6435875831d1https://doi.org/10.24963/ijcai.2024/90
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1OV-Uni3DETR: Towards Unified Open-Vocabulary 3D Object Detection via Cycle-Modality Propagation2024 · 1 citations
  2. 2A Unified Framework for 3D Scene Understanding2024 · 1 citations
  3. 3Lowis3D: Language-Driven Open-World Instance-Level 3D Scene Understanding2024 · 26 citations
  4. 4Open-Vocabulary SAM3D: Towards Training-free Open-Vocabulary 3D Scene Understanding2024
  5. 53D Annotation-Free Learning by Distilling 2D Open-Vocabulary Segmentation Models for Autonomous Driving2024