PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 18, 20244 citationsOpen Access

Audio-Journey: Open Domain Latent Diffusion Based Text-To-Audio Generation

View Full Paper
JMJackson MichaelsJLJuncheng B LiLYLaura Yao

Key Points

Key points are not available for this paper at this time.

Abstract

Despite recent progress, machine learning (ML) models for open-domain audio generation need to catch up to generative models for image, text, speech, and music. The lack of massive open-domain audio datasets is the main reason for this performance gap; we overcome this challenge through a novel data augmentation approach. We leverage state-of-the-art (SOTA) Large Language Models (LLMs) to enrich captions in the weakly-labeled audio dataset. We then use a SOTA video-captioning model to generate captions for the videos from which the audio data originated, and we again use LLMs to merge the audio and video captions to form a rich, large-scale dataset. We experimentally evaluate the quality of our audio-visual captions, showing a 12.5% gain in semantic score over baselines. Using our augmented dataset, we train a Latent Diffusion Model to generate in an encodec encoding latent space. Our model is novel in the current SOTA audio generation landscape due to our generation space, text encoder, noise schedule, and attention mechanism. Together, these innovations provide competitive open-domain audio generation. The samples, models, and implementation will be at https://audiojourney.github.io.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Michaels et al. (2024) studied this question.

synapsesocial.com/papers/68e7398bb6db6435876b2c62https://doi.org/10.1109/icassp48485.2024.10448220
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners2024 · 1 citations
  2. 2Improving Audio Generation with Visual Enhanced Caption2024
  3. 3Generating Moving 3D Soundscapes with Latent Diffusion Models2025
  4. 4AudioLCM: Text-to-Audio Generation with Latent Consistency Models2024 · 1 citations
  5. 5Video-to-Audio Generation with Hidden Alignment2024 · 1 citations