PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 19, 2026ISPRS International Journal of Geo-Information1 citationsOpen Access

VGGT-Geo: Probabilistic Geometric Fusion of Visual Geometry Grounded Transformer Priors for Robust Dense Indoor SLAM

View Full Paper
KQKai QinJLJing LiSZSisi Zlatanova

Key Points

  • The research aims to develop a robust SLAM system that combines geometric and learning-based approaches to improve 3D perception in challenging environments.
  • Developed VGGT-Geo integrating generative priors with geometric optimization.
  • Implemented a Probabilistic Geometric Fusion framework.
  • Created a Confidence-Aware Optimization for feature extraction and confidence prediction.
  • Used Multi-Modal Constraint Closure to enhance feature fusion in texture-less areas.
  • VGGT-Geo achieved a superior Absolute Trajectory Error of 4-5 cm.
  • Demonstrated a Relative Rotation Error of 0.79° in complex environments.
  • Outperformed existing state-of-the-art methods by about 50% in trajectory accuracy.

Abstract

With the rapid evolution of Digital Twins and Embodied AI, achieving fast, dense, and high-precision 3D perception in unknown environments has become paramount. However, existing Visual SLAM paradigms face a critical dilemma: geometry-based methods often fail in texture-less areas due to feature scarcity, while learning-based approaches frequently suffer from scale drift and unphysical deformations. To bridge this gap, we propose VGGT-Geo, a novel SLAM system that synergizes generative priors from Large Foundation Models with multi-modal geometric optimization. Distinguishing itself from simple cascaded architectures, we construct a Probabilistic Geometric Fusion framework, consisting of (1) Generative Warm-start, leveraging the holistic scene understanding capabilities of the VGGT, (2) Confidence-Aware Optimization to extract dense features via DINOv3 and predict their confidence map, and (3) a Multi-Modal Constraint Closure that fuses point-line features and metric depth priors to constrain rotational Degrees of Freedom in Manhattan Worlds. We conducted systematic evaluations on TUM, Replica, Tanks and Temples, and a challenging self-collected dataset featuring extreme lighting and texture-less walls. Experimental results demonstrate that VGGT-Geo exhibits superior robustness and accuracy in unseen environments. On our most challenging dataset, it achieves an Absolute Trajectory Error of 4-5 cm and a Relative Rotation Error of 0.79°, outperforming current state-of-the-art methods by approximately 50% in trajectory accuracy. This study validates that synergizing the intuition of Large Foundation Models with geometric rigor is a viable path toward next-generation robust SLAM.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Qin et al. (2026) studied this question.

synapsesocial.com/papers/6996a8efecb39a600b3f0297https://doi.org/10.3390/ijgi15020085
Ask AI
Helpful
Bookmark
Share
View Full Paper