PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 29, 20241 citationsOpen Access

DriveGenVLM: Real-world Video Generation for Vision Language Model based Autonomous Driving

View Full Paper
YFYongjie FuAJAnmol JainXDXuan Di

Key Points

Key points are not available for this paper at this time.

Abstract

The advancement of autonomous driving technologies necessitates increasingly sophisticated methods for understanding and predicting real-world scenarios. Vision language models (VLMs) are emerging as revolutionary tools with significant potential to influence autonomous driving. In this paper, we propose the DriveGenVLM framework to generate driving videos and use VLMs to understand them. To achieve this, we employ a video generation framework grounded in denoising diffusion probabilistic models (DDPM) aimed at predicting real-world video sequences. We then explore the adequacy of our generated videos for use in VLMs by employing a pre-trained model known as Efficient In-context Learning on Egocentric Videos (EILEV). The diffusion model is trained with the Waymo open dataset and evaluated using the Fr\'echet Video Distance (FVD) score to ensure the quality and realism of the generated videos. Corresponding narrations are provided by EILEV for these generated videos, which may be beneficial in the autonomous driving domain. These narrations can enhance traffic scene understanding, aid in navigation, and improve planning capabilities. The integration of video generation with VLMs in the DriveGenVLM framework represents a significant step forward in leveraging advanced AI models to address complex challenges in autonomous driving.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Fu et al. (2024) studied this question.

synapsesocial.com/papers/68e5a80fb6db64358754239bhttps://doi.org/10.48550/arxiv.2408.16647
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation2024 · 4 citations
  2. 2DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models2024 · 23 citations
  3. 3ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving2025
  4. 4Less is More: Lean yet Powerful Vision-Language Model for Autonomous Driving2025
  5. 5EVGen: Trajectory-conditioned forward-view video generation under minimal visual observations2026